Hiring.Camp

AI Infrastructure Software Engineer — CosmosLab

Nvidia

·

2 weeks ago

Location
Beijing, Beijing,CN, CN · Shanghai, Shanghai,CN, CN · Shenzhen, Guangdong,CN, CN
Type
Full-time
Department
Engineering
Education
Bachelor
Source
Eightfold

Description

Welcome to NVCareers!

The place to find available career opportunities at NVIDIA for you and people you know.

  • If you're interested in growing your career at NVIDIA, click the Apply button and follow the instructions on the next page.
  • For more information about referring a friend to this job, please review this guide.
  • If you would like to share this job to your social network, follow the steps here.

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

Are you excited to explore new frontiers in AI? Join NVIDIA’s Cosmos Lab Infra team and take part in the innovation building the training infrastructure that supports our Physical AI world foundation models. Here is your opportunity to design, assemble, and improve the infrastructure for large-scale AI training, spanning pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL) post-training. Embark on a journey where your work will be essential to influencing the future of AI!

What you'll be doing:

  • Create and implement the training infrastructure spanning pre-training, SFT, and RL post-training for Physical AI world foundation models. The work involves the framework and a comprehensive control plane across clusters to coordinate workloads efficiently.
  • Develop and improve the pre-training and SFT pipelines — large-scale data loading, distributed training, and checkpointing — to achieve high throughput and scalability.
  • Develop and improve the inference and evaluation stack, including the inference engine, inference/generation pipelines (which also support RL rollout), and evaluation pipelines. Use methods like continuous batching and KV-cache management to achieve high throughput and low latency.
  • Build and improve the effective interaction and data flow among the RL system's roles (policy, rollout, reward, simulation) while investigating system-level optimization opportunities.
  • Integrate and orchestrate simulation and robotics environments as RL environments — driving the simulation↔rollout↔training loop at scale.
  • Build and refine the distributed training backend — sharding/parallelism, mixed precision, activation checkpointing, and memory/throughput optimization across many GPUs.
  • Improve the efficiency, scalability, and resiliency of training and RL workloads — focusing on fault tolerance, fast/elastic restart, and throughput optimization under preemption and hardware failure.
  • Define meaningful, actionable reliability and efficiency metrics to track and improve system reliability.
  • Root cause, triage, and resolve failures from the application level down to the framework, GPU, and network/hardware level.

What we need to see:

  • 5+ years developing software infrastructure for large-scale AI or distributed systems.
  • Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).
  • Strong debugging and triage skills across the stack — from AI application down to GPU/hardware behavior.
  • Proven track record building and scaling large-scale distributed systems, ideally distributed training or inference.
  • Hands-on experience with AI training and/or inference infrastructure — RL/post-training, training frameworks, or inference serving.
  • Proficiency in Python (plus scripting), and solid software engineering practices: testing, defensive programming, version control, and CI.
  • Excellent communication and collaboration skills; intellectual curiosity, problem-solving, and willingness.

Ways to stand out from the crowd:

  • Experience building RL / post-training infrastructure — PPO/GRPO/DPO pipelines, rollout engines, and asynchronous RL.
  • Background with building large-scale, production-grade pre-training / SFT infrastructure.
  • Experience integrating simulation / robotics environments into training or RL loops — including vectorized environments and sim-to-real workflows.
  • Comprehensive knowledge of DL framework internals — PyTorch (FSDP/DTensor) and Megatron or equivalent experience, distributed training, and related optimization techniques.
  • Proficiency in C/C++/CUDA for performance-critical components and custom kernels.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Skills

PythonPyTorch

Similar Jobs

30

AI Infrastructure Software Engineer

Qualcomm · Shanghai, Shanghai,CN, CN

5 days ago

AI Infrastructure Software Engineer — CosmosLab

Nvidia · China, Beijing +2

1 week ago

AI Infrastructure Software Engineer

Qualcomm · Shanghai, Shanghai,CN, CN

1 month ago

HPC / AI Software Infrastructure Lead (E)

KLA · USA-MI-Ann Arbor-KLA, United States of America · onsite

1 month ago

HPC / AI Software Infrastructure Lead (E)

Kla · USA-MI-Ann Arbor-KLA, United States of America

1 month ago

Senior DGX Cloud AI Infrastructure Software Engineer

Nvidia · China, Shanghai

2 months ago

Senior DGX Cloud AI Infrastructure Software Engineer

Nvidia · Shanghai, Shanghai,CN, CN

2 months ago

Senior AI Infrastructure Software Engineer - DGX Cloud

Nvidia · US, CA, Santa Clara, United States of America +3 · Remote

2 months ago

Senior AI Infrastructure Software Engineer - DGX Cloud

Nvidia · Santa Clara, CA,US, US +3 · Remote

2 months ago

Sr. AI Infrastructure Software Engineer

KLA · USA-CA-Milpitas-KLA, United States of America · Remote

4 months ago

Senior DGX Cloud AI Infrastructure Software Engineer

Nvidia · US, CA, Santa Clara, United States of America +4 · Remote

5 months ago

Senior DGX Cloud AI Infrastructure Software Engineer

Nvidia · Remote

5 months ago

Senior AI Infrastructure Software Engineer

Nvidia · China, Shanghai · Remote

6 months ago

Senior AI Infrastructure Software Engineer

Nvidia · Remote

6 months ago

Senior DGX Cloud AI Infrastructure Software Engineer

Nvidia · Shanghai, Shanghai,CN, CN · Remote

8 months ago

Sr. Software Engineer, AI Infrastructure

LinkedIn · Sunnyvale, CA, United States · Hybrid

4 days ago

Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation

Ambient.Ai · Redwood City · Hybrid

1 week ago

Manager, Software Engineering - AI Infrastructure

Nvidia · Israel, Tel Aviv +1

1 week ago

Manager, Software Engineering - AI Infrastructure

Nvidia · Tel Aviv-Yafo, Tel Aviv District,IL, IL +1

1 week ago

Backend Infrastructure & Agentic AI Software Engineer

Booz Allen Hamilton · USA, DC, Washington (1301 K St), United States of America

1 week ago

Federal Account Executive - Cybersecurity, AI, or Infrastructure Software Sales

Talent Search PRO · Onsite

2 weeks ago

Senior System Software Engineer, AI Infrastructure

Nvidia · US, CA, Santa Clara, United States of America

3 weeks ago

Senior System Software Engineer, AI Infrastructure

Nvidia · Santa Clara, CA,US, US

3 weeks ago

Staff Software Engineer (AI Infrastructure/Python)

NBCUniversal · New York, NEW YORK, United States · Remote

1 month ago

Senior Software Engineer, AI Infrastructure

Thealleninstitute · Seattle, WA +1 · Onsite

1 month ago

Principal Software Developer, AI Infrastructure

Oracle · Austin, TX, United States

2 months ago

Software Engineer, AI Infrastructure

Glean · San Francisco, California, United States +1

2 months ago

Senior Software Engineer - AI Infrastructure

Oracle · Austin, TX, United States

2 months ago

Staff Software Engineer, Infrastructure – AI Platform

Addepar1 · Pune, India +1

2 months ago

AI/ML Infrastructure Software Development Engineer

Booz Allen Hamilton · USA, DC, Washington (1301 K St), United States of America

3 months ago