Hiring.Camp

AI Infrastructure and Frameworks Intern, Cosmos Lab - 2027

Nvidia

·

Today

Location
Beijing, Beijing,CN, CN · Shanghai, Shanghai,CN, CN · Shenzhen, Guangdong,CN, CN
Type
Internship
Department
IT
Seniority
Internship
Education
PhD
Source
Eightfold

Description

Join NVIDIA’s Cosmos Lab Infrastructure team to develop training and post-training systems for advanced Physical AI models, including world foundation models and robot policies. Our infrastructure connects training, inference, and evaluation with simulation and real-world robot interaction. You will work with a mentor on a focused project scoped to your experience and internship duration, implementing and evaluating systems improvements on real AI workloads using NVIDIA’s GPU infrastructure.

What you’ll be doing:

  • Develop and optimize training infrastructure for advanced Physical AI world models, supporting pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL). Explore distributed parallelism, sharding, low-precision training, compute–communication overlap, and numerical consistency and efficient weight synchronization between training and inference.
  • Build Physical AI post-training and RL infrastructure supporting advanced training algorithms. Connect simulation or, where applicable, real-robot interaction with experience collection, rollout inference, reward computation, training, and evaluation. Optimize these workflows through partitioning, pipelining, data transfer, and synchronization across synchronous, asynchronous, or disaggregated execution.
  • Improve efficiency and scalability across training, inference, simulation, and evaluation through scheduling, placement, dynamic resource allocation, and load balancing, supporting heterogeneous resources, elasticity, and fault recovery.
  • Analyze and optimize system performance, working with researchers to investigate, support, and compare emerging Physical AI models, training workflows, and algorithms from a systems perspective. Use profiling, benchmarking, and performance modeling to identify bottlenecks and measure throughput, latency, GPU utilization, and policy freshness. Share findings through tested code, documentation, and technical presentations, and contribute to research publications where appropriate.

What we need to see:

  • Pursuing a Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • Strong Python and debugging skills, with systems fundamentals in concurrency, distributed execution, memory management, or data movement.
  • Practical experience in at least one area: training infrastructure, RL infrastructure, simulation or robotics integration, or inference infrastructure. Coursework, research, open-source projects, and internships all count.
  • Strong analytical and communication skills, curiosity, and a willingness to learn.
  • Experience in every listed area, prior access to large GPU clusters, and model architecture or learning algorithm research are not required.

Ways to stand out from the crowd:

  • Experience optimizing training infrastructure, including distributed parallelism, low-precision training, GPU memory efficiency, or compute–communication overlap.
  • Experience optimizing scheduling, placement, resource allocation, or data transfer across training, rollout, simulation, and evaluation.
  • Experience extending RL pipelines, integrating simulation environments or robot interfaces, or optimizing inference; GPU profiling, C++/CUDA development, and open-source contributions or research in ML systems are also valued.

Skills

Python

Similar Jobs

30

AI Infrastructure and Frameworks Intern, Cosmos Lab - 2027

Nvidia · China, Beijing +2

Today

Staff Product Manager, AI Infrastructure and Compliance

Boxinc · Redwood City, CA +1

6 days ago

Research Computing Engineer (AI Infrastructure and HPC)

Psu · Penn State University Park, United States of America · Remote, Hybrid, Onsite

1 week ago

Research Computing Engineer (AI Infrastructure and HPC)

Psu · Penn State University Park, United States of America · Remote, Hybrid, Onsite

1 week ago

Solutions Engineer-Hyperscale and AI Infrastructure

Cisco · USA-RICHARDSON, United States of America +1 · Remote

1 week ago

Senior Manager, AI Infrastructure and Networking

Verizon Communications · One Verizon Way, Basking Ridge, NJ (NJ0533), United States of America +8

2 weeks ago

Principal Product Manager, AI Infrastructure and Orchestration

DataRobot delivers AI that maximizes · Remote WA, United States of America +3 · Remote

1 month ago

Product Manager - Connectors and AI Infrastructure

SafeBreach · Tel Aviv-Yafo, Israel +1 · Hybrid

1 month ago

Senior HPC Support Engineer - Ethernet and AI Infrastructure

Nvidia · TX,US, US · Remote

2 months ago

Senior HPC Support Engineer - Ethernet and AI Infrastructure

Nvidia · US, TX, Remote, United States of America · Remote

2 months ago

Senior HPC Support Engineer - Ethernet and AI Infrastructure

Nvidia · TN,US, US · Remote

2 months ago

Senior HPC Support Engineer - Ethernet and AI Infrastructure

Nvidia · US, TN, Remote, United States of America · Remote

2 months ago

Engineering Manager, Platform and AI Infrastructure

Anrok · San Francisco +3 · Remote

3 months ago

Principal Product Manager Hardware, Cloud Infrastructure and AI Networking

Hewlett Packard Enterprise · All, North Carolina, United States of America

3 months ago

Principal Product Manager Hardware, Cloud Infrastructure and AI Networking

Hewlett Packard Enterprise · All, North Carolina, United States of America

3 months ago

Principal Product Manager Hardware, Cloud Infrastructure and AI Networking

Hewlett Packard Enterprise · All, North Carolina, United States of America

3 months ago

Senior Faculty Lead, Digital Pathology Infrastructure and AI Enablement

Harvard Medical Faculty Physicians · BIDMC - Main Campus, United States of America

4 months ago

Principal Product Manager Hardware, Cloud Infrastructure and AI Networking

Hewlett Packard Enterprise (HP) · Sunnyvale, California, United States of America · Onsite

6 months ago

Principal Product Manager Hardware, Cloud Infrastructure and AI Networking

Hewlett Packard Enterprise (HP) · Sunnyvale, California, United States of America · Onsite

6 months ago

Principal Product Manager Hardware, Cloud Infrastructure and AI Networking

Hewlett Packard Enterprise (HP) · Sunnyvale, California, United States of America · Onsite

6 months ago

Director, Public Relations – AI and Cloud Infrastructure (Location: Mid-West Area)

Oracle · IL, United States, US

Today

Director, Public Relations – AI and Cloud Infrastructure (Location: Southwest Area)

Oracle · AZ, United States, US

2 days ago

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AI · Bangalore India +1 · Remote

6 days ago

Intern - AI Systems and Infrastructure Engineering

Micron Technology · Austin, TX,US, US

2 weeks ago

Intern - AI Systems and Infrastructure Engineering

Micron · Austin, TX, United States of America

2 weeks ago

Director of AI Platforms and Infrastructure 2026- US

Aimpoint Digital · Atlanta, GA, US · Remote

1 month ago

AI-Enabled Infrastructure and Systems Engineer

BlackRock · NY7 - 50 Hudson Yards, New York, United States of America

2 months ago

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together · San Francisco +1

3 months ago

AI Compute and Infrastructure Counsel (Legal)

Reflectionai · San Francisco +1 · Onsite

4 months ago

Software Engineer, AI Training and Infrastructure

Skildai Careers · Pittsburgh, San Francisco, Bengaluru +1 · Remote

1+ year ago