- Salary
- $150k – $450k/yr
- Location
- Sunnyvale, CA
- Workplace
- Onsite
- Type
- Full-time
- Department
- Engineering
- Experience
- 5+ years
- Source
- Lever
Description
The Role
We’re looking for an RL infrastructure engineer to help extend and scale our end-to-end RL training systems. You’ll work side-by-side with world-class researchers and engineers to:
- Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
- Integrate rollout generation, reward computation, trajectory processing, policy updates, and weight synchronization
- Build robust config + launch systems across multi-node, multi-GPU clusters
- Own experiment tracking, metrics logging, and job monitoring for external visibility
- Improve training system reliability, maintainability, and performance
Hands-on experience developing end-to-end RL infrastructure is required. Strong infrastructure and systems experience is what we value most.