- Salary
- $150k – $450k/yr
- Location
- Sunnyvale, CA
- Workplace
- Onsite
- Type
- Full-time
- Department
- Engineering
- Experience
- 5+ years
- Source
- Lever
Description
The Role
We’re looking for a distributed ML infrastructure engineer to help extend and scale our training systems. You’ll work side-by-side with world-class researchers and engineers to:
- Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
- Implement distributed optimizers from mathematical specs
- Build robust config + launch systems across multi-node, multi-GPU clusters
- Own experiment tracking, metrics logging, and job monitoring for external visibility
- Improve training system reliability, maintainability, and performance
Much of the work will support large-scale pre-training, and pre-training experience is required. Strong infrastructure and systems experience is what we value most.