- Location
- Redmond, WA,US, US · Santa Clara, CA,US, US
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Closing date
- Today
- Source
- Eightfold
Description
Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang) Own model export pipelines (ModelOpt, Megatron-LM <-> HuggingFace), ensuring quantized checkpoints serialize correctly for downstream serving Build prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimization Develop data analysis tooling and visualizations for numerics debugging Improve developer productivity across the team: CI, build systems, training infrastructure, pipeline friction Participate in code reviews and incorporate feedback Proficient in Python; familiarity with C++ Experience with ML accelerators with a basic understanding of how certain ML layers affect execution time Experience reading, modifying, or contributing to a large open-source codebase Experience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang) or Triton kernel development