Hiring.Camp

AI Infrastructure Software Engineer — CosmosLab

Nvidia

·

Today

Location
China, Shanghai · China, Beijing · China, Shenzhen
Type
Full-time
Department
Engineering
Education
Bachelor
Source
Workday

Description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

Are you excited to explore new frontiers in AI? Join NVIDIA’s Cosmos Lab Infra team and take part in the innovation building the training infrastructure that supports our Physical AI world foundation models. Here is your opportunity to design, assemble, and improve the infrastructure for large-scale AI training, spanning pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL) post-training. Embark on a journey where your work will be essential to influencing the future of AI!

What you'll be doing:

  • Create and implement the training infrastructure spanning pre-training, SFT, and RL post-training for Physical AI world foundation models. The work involves the framework and a comprehensive control plane across clusters to coordinate workloads efficiently.

  • Develop and improve the pre-training and SFT pipelines — large-scale data loading, distributed training, and checkpointing — to achieve high throughput and scalability.

  • Develop and improve the inference and evaluation stack, including the inference engine, inference/generation pipelines (which also support RL rollout), and evaluation pipelines. Use methods like continuous batching and KV-cache management to achieve high throughput and low latency.

  • Build and improve the effective interaction and data flow among the RL system's roles (policy, rollout, reward, simulation) while investigating system-level optimization opportunities.

  • Integrate and orchestrate simulation and robotics environments as RL environments — driving the simulation↔rollout↔training loop at scale.

  • Build and refine the distributed training backend — sharding/parallelism, mixed precision, activation checkpointing, and memory/throughput optimization across many GPUs.

  • Improve the efficiency, scalability, and resiliency of training and RL workloads — focusing on fault tolerance, fast/elastic restart, and throughput optimization under preemption and hardware failure.

  • Define meaningful, actionable reliability and efficiency metrics to track and improve system reliability.

  • Root cause, triage, and resolve failures from the application level down to the framework, GPU, and network/hardware level.

What we need to see:

  • 5+ years developing software infrastructure for large-scale AI or distributed systems.

  • Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).

  • Strong debugging and triage skills across the stack — from AI application down to GPU/hardware behavior.

  • Proven track record building and scaling large-scale distributed systems, ideally distributed training or inference.

  • Hands-on experience with AI training and/or inference infrastructure — RL/post-training, training frameworks, or inference serving.

  • Proficiency in Python (plus scripting), and solid software engineering practices: testing, defensive programming, version control, and CI.

  • Excellent communication and collaboration skills; intellectual curiosity, problem-solving, and willingness.

Ways to stand out from the crowd:

  • Experience building RL / post-training infrastructure — PPO/GRPO/DPO pipelines, rollout engines, and asynchronous RL.

  • Background with building large-scale, production-grade pre-training / SFT infrastructure.

  • Experience integrating simulation / robotics environments into training or RL loops — including vectorized environments and sim-to-real workflows.

  • Comprehensive knowledge of DL framework internals — PyTorch (FSDP/DTensor) and Megatron or equivalent experience, distributed training, and related optimization techniques.

  • Proficiency in C/C++/CUDA for performance-critical components and custom kernels.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/ 

Skills

PythonPyTorch

Similar Jobs

30

#AI Infrastructure Software Engineer

Qualcomm · San Diego, CA,US, US

1 week ago

AI Infrastructure Software Engineer — CosmosLab

Nvidia · Beijing, Beijing,CN, CN +2

1 month ago

HPC / AI Software Infrastructure Lead (E)

KLA · USA-MI-Ann Arbor-KLA, United States of America · onsite

1 month ago

Sr. AI Infrastructure Software Engineer

KLA · USA-CA-Milpitas-KLA, United States of America · Remote

5 months ago

Senior DGX Cloud AI Infrastructure Software Engineer

Nvidia · Remote

6 months ago

Senior Web Software Engineer, AI Infrastructure

Nvidia · China, Shanghai

1 week ago

Senior Web Software Engineer, AI Infrastructure

Nvidia · Shanghai, Shanghai,CN, CN

1 week ago

Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation

Ambient.Ai · Redwood City · Hybrid

4 weeks ago

Senior Software Engineer, AI Infrastructure

Thealleninstitute · Seattle, WA +1 · Onsite

2 months ago

Principal Software Developer, AI Infrastructure

Oracle · Austin, TX, United States

2 months ago

Staff Software Engineer, Infrastructure – AI Platform

Addepar1 · Pune, India +1

3 months ago

Senior Software Engineer, Infrastructure AI

Airwallex · US - Seattle · Remote

3 months ago

Senior Software Engineer – AI Infrastructure

Kraken · United Kingdom +14 · Remote

5 months ago

Software Engineer, AI Infrastructure

Fireworksai · New York, NY; San Mateo, CA +1 · Remote

9 months ago

Software Developer 5 , AI Infrastructure

Oracle · Seattle, WA, United States, US

1 month ago

Staff Software Engineer, Infrastructure (Data & AI)

Airwallex · SG - Singapore · Onsite

1 month ago

Senior Software Engineer, AI Speed Infrastructure

Nvidia · Santa Clara, CA,US, US +1

2 months ago

Senior Software Engineer, AI Speed Infrastructure

Nvidia · US, CA, Santa Clara, United States of America +1

2 months ago

Backend Infrastructure & Agentic AI Platforms Software Development Engineer, Senior

Booz Allen Hamilton · USA, DC, Washington (1301 K St), United States of America +1

2 months ago

Software Engineer, Agentic AI Infrastructure

Anrok · San Francisco +2 · Hybrid

2 months ago

Software Engineer, Cortex AI Infrastructure

Snowflake · US-CA-Menlo Park

2 months ago

Senior Software Engineer - Verification AI Infrastructure

Nvidia · Israel, Tel Hai

3 months ago

Software Engineer, AI Compute Infrastructure

Heygen · Los Angeles, Palo Alto, San Francisco, Toronto, Singapore +4 · Remote

8 months ago

Staff Infrastructure Software Engineer, Enterprise AI

Scaleai · New York, NY; San Francisco, CA +1

11 months ago

Senior Software Engineer (Python Backend & AI Infrastructure)

credium GmbH · Augsburg, Bayern +1

1 week ago

Staff Software Development Engineer - Enterprise AI Infrastructure - #4898

Grailbio · Menlo Park, CA +1 · Hybrid

3 weeks ago

Software Engineer, DGX Cloud AI Infrastructure

Nvidia · US, CA, Santa Clara, United States of America +4 · Remote

4 weeks ago

Senior Software Engineer, DGX Cloud AI Infrastructure

Nvidia · Santa Clara, CA,US, US +4 · Remote

2 months ago

Senior Software Engineer, DGX Cloud AI Infrastructure

Nvidia · US, CA, Santa Clara, United States of America +4 · Remote

2 months ago

Software Engineer, DGX Cloud AI Infrastructure

Nvidia · Santa Clara, CA,US, US +4 · Remote

2 months ago