Hiring.Camp

Senior Distributed Systems Engineer

Institute of Foundation Models

·

Mar 3, 2026

Salary
$200k – $400k/yr
Location
Sunnyvale, CA
Workplace
Onsite
Department
Engineering
Seniority
Senior
Source
Lever

Description

About the Institute of Foundation Models
The Institute of Foundation Models (IFM) designs and operates ultra-scale GPU supercomputing systems to train next-generation foundation models. We believe performance, fault tolerance, and scalability are co-designed across model architecture, communication systems, runtime, and hardware topology.
This role sits at the core of that effort — driving communication performance, distributed reliability, and cross-layer optimization for large-scale training workloads.

The Mission
We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including hybrid parallelism and Mixture-of-Experts (MoE) workloads.
This is not a network operations role. This is a systems-level engineering position focused on performance engineering, distributed debugging, and communication-runtime co-design.
·       Design and optimize expert-parallel and hybrid-parallel communication patterns
·       Drive high-performance hierarchical collectives for MoE workloads
·       Co-design runtime orchestration with communication topology awareness
·       Reduce tail latency and improve determinism across thousands of GPUs
·       Architect fault-tolerant distributed execution under real-world cluster failures
Core Technical Scope
·       Communication-compute overlap and topology-aware collective optimization
·       Deep debugging of NCCL, RDMA, and custom communication layers
·       Hybrid expert parallel strategies in modern large-scale MoE systems
·       Elastic and resilient distributed job orchestration concepts
·       Congestion analysis and routing optimization across InfiniBand/RoCE fabrics
·       Microbenchmarking and performance modeling for communication-heavy workloads
Expected Technical Depth
·       Hybrid expert parallel communication for Mixture-of-Experts training
·       Scaling behavior under network pressure
·       Distributed orchestration for elastic, large-scale training
·       Fault detection and recovery in distributed GPU workloads
·       Cross-layer bottlenecks: GPU ↔ NIC ↔ PCIe ↔ NVSwitch ↔ Fabric ↔ Scheduler
Required Background
·       Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth)
·       Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA
·       Deep familiarity with NCCL and/or UCX internals
·       Strong systems programming ability (C/C++, Rust, or Go)
·       Strong familiarity with modern model training frameworks such as PyTorch
·       Ability to troubleshoot and profile training performance issues related to communication bottlenecks
·       Ability to translate research ideas into production-grade optimizations
·       Experience debugging distributed hangs, desynchronization, and performance regressions
What We Mean by "Hardcore"
·       You can explain why an communication degrades at scale and how to fix it
·       You have improved real cluster throughput via communication redesign
·       You can trace a distributed hang across ranks and identify the root cause
·       You are comfortable working at the boundary between hardware and runtime
Application Requirements
·       Include a link to your GitHub (required)
·       Provide links to relevant distributed systems, HPC, or large-scale training projects
·       Include a list of publications and/or public technical reports (if applicable)
·       Describe the hardest distributed debugging problem you solved
·       Include measurable performance improvements you have delivered
Academic Qualifications
Master’s, or Bachelor’s + 1 year of relevant experience.

Skills

RustPyTorchGitHub

Similar Jobs

30

Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure

Nvidia · US, WA, Remote, United States of America +5 · Remote

3 days ago

Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure

Nvidia · WA,US, US +5 · Remote

3 days ago

Senior Software Engineer — Distributed Compute / Spark Systems

Granica · Bay Area Office · Hybrid

1 week ago

Senior Distributed Systems Engineer

Censys · Remote (US/Canada) · Remote

1 week ago

Staff/Senior Distributed Systems Engineer

Helios Intelligence Platforms · New York City · Onsite

1 week ago

Senior Lead AI Engineer (Gen AI Platform Services: Distributed Systems)

Capitalone · Cambridge, MA, United States of America +4

2 weeks ago

Sr Software Engineer - Cloud Platform—Kubernetes, Hyperscalers & Distributed Systems

ServiceNow · Hyderabad, India · Hybrid

3 weeks ago

Senior Software Systems Architect_Edge Computing & Distributed Platforms

Thales · Esch-Belval, Luxembourg

3 weeks ago

Senior Software Engineer - Distributed Systems

Workday · Ireland, Dublin

1 month ago

Sr. Software Engineer (Distributed Systems)

Workday · Ireland, Dublin

1 month ago

Senior Distributed Systems Engineer

Ionq · Santa Clara, California, United States · Remote, Onsite

1 month ago

Senior Distributed Systems Engineer

Cloudflare · In-Office +1 · Onsite

1 month ago

Senior Staff Distributed Systems Engineer

Ionq · Santa Clara, California, United States · Remote, Onsite

1 month ago

Senior Backend Software Engineer - Distributed Systems

HavocAI · Remote · Remote

1 month ago

Distributed Systems - Senior Staff Systems Architect

Northrop Grumman · Baltimore, MD,US, US

1 month ago

Distributed Systems - Senior Staff Systems Architect

Northrop Grumman · MDLI04, United States of America

1 month ago

Sr. Staff Software Development Engineer - Go/Distributed Systems

Zscaler · Hyderabad, IND · Hybrid

2 months ago

Senior Distributed Systems Engineer - Platform Engineering

RTB House · Poland

2 months ago

Senior Software Engineer, Engine & Distributed Systems

StackAI · Poland - Warsaw Office · Onsite

2 months ago

Senior Software Engineer, Engine & Distributed Systems

StackAI · SF Office - 171 2nd, 4th floor +1 · Hybrid

2 months ago

Senior Software Engineer (C#, Distributed systems)

Omnissa · India-Bangalore-Office-Kalyani Vista

2 months ago

Senior Software Engineer, Infrastructure Automation and Distributed Systems

Nvidia · SC,US, US +4 · Remote

2 months ago

Senior Software Engineer, Infrastructure Automation and Distributed Systems

Nvidia · US, SC, Remote, United States of America +4 · Remote

2 months ago

Sr. Software Engineer (Distributed Systems)

Workday · USA, GA, Atlanta, United States of America

2 months ago

Senior Software Engineer — Platform & Distributed Systems (XTM Foundation)

Filigran · France · Remote

2 months ago

Senior Engineer – Java Bigdata Kafka Distributed Systems – Assistant Vice President

Citi Bank · IT BUILDING, RAMANUJAN IT SEZ,, India +1 · Hybrid

3 months ago

Senior Lead Software Engineer, Distributed Systems (Golang + Python on Kubernetes)

Capitalone · San Francisco, CA, United States of America +4

3 months ago

Senior Software Engineer Cloud Architecture & Distributed Systems

WSP · Noida, Uttar Pradesh, India

3 months ago

Senior Engineer – Java Bigdata Kafka Distributed Systems – Assistant Vice President

citibank · Chennai, TN,IN, IN

3 months ago

Senior Lead Software Engineer, Distributed Systems (Golang + Python on Kubernetes)

Capitalone · San Francisco, CA, United States of America +4

3 months ago