- Salary
- $200k – $400k
- Location
- San Francisco
- Workplace
- Remote, Onsite
- Type
- Full-time
- Department
- Engineering
- Closing date
- Today
- Source
- CareersPage
Description
Machine Learning Engineer - Infrastructure
Company: Causal Labs
Location: San Francisco, CA (South Park office, in person 5 days per week; relocation provided)
Compensation: $200,000 - $400,000 + highly competitive early-stage equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers; can sponsor visas
About Causal Labs
Causal Labs is pursuing general causal intelligence: AI that can predict the future and identify the actions that change it. It is building a Large Physics foundation Model (LPM), because domains governed by physics have inherent cause-and-effect structure that visual or textual data lacks. Its starting domain is weather, the most observed physical system on earth, with rapid ground-truth feedback and data volumes that dwarf LLM training sets.
The founders come from Cruise, Google Research and Meta. The company is about 10 people in San Francisco, growing to around 35 this year, and is backed by Kindred Ventures, Refactor and BoxGroup.
The Role
Causal Labs is hiring infrastructure engineers to tackle the unsolved training and inference challenges of a Large Physics foundation Model. The work demands deep expertise in standing up distributed training clusters and optimizing performance for large models. If you have built large-scale ML infrastructure for language, vision, robotics or biology models and want to bet on a counterintuitive technical thesis, this is the role.
What You Will Do
- Design, deploy and maintain large distributed ML training and inference clusters.
- Build efficient, scalable end-to-end pipelines for petabyte-scale datasets and model training across the ML lifecycle.
- Research and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scales.
- Analyze, profile and debug low-level GPU operations to optimize performance.
- Bring new ideas from current research into the stack.
What You Bring
- 2-10 years building large-scale ML infrastructure for core foundation models (not fine-tuning)
- Deep expertise optimizing large-scale training and inference workloads
- Proficiency with distributed training frameworks (FSDP, DeepSpeed)
- Knowledge of cloud platforms (GCP, AWS or Azure) and containers/orchestration (Kubernetes, Docker)
- Experience with distributed task management and scalable model serving architectures
- Strong grasp of monitoring, logging and observability for ML systems
- Ability to work in person in San Francisco 5 days a week
Nice to Have
- Experience at a science or physical AI company (self-driving, robotics, biology, climate/weather)
- Generalist experience across the ML lifecycle
- Low-level GPU performance optimization and debugging (CUDA, JAX)
Interview Process
Initial call (30 min), technical screen, onsite day.
Tech Stack
FSDP, DeepSpeed, NVIDIA GPUs, Python, C++, Linux, Kubernetes, Docker, GCP/AWS/Azure