Hiring.Camp

Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems

Nvidia

·

Jun 26, 2026

Location
Santa Clara, CA,US, US
Type
Full-time
Department
Engineering
Seniority
Director
Source
Eightfold

Description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

NVIDIA is seeking a strong technology leader for our Engineering Operations and Site Reliability Engineering for our next-generation datacenter server systems. This role sits at the intersection of execution, reliability, automation, and large-scale system operations, where we keep NVIDIA’s rack-scale systems healthy, observable, and highly available for internal engineering users. These systems bring together the full power of NVIDIA CPUs, GPUs, NVLink, InfiniBand/Spectrum-X networking, cluster management technologies, and our optimized AI/HPC software stack. We enable fast product development by ensuring large internal racks, clusters, and lab infrastructure are reliable, well-instrumented, and operated with scalable engineering practices. This is a technical leadership role focused on execution excellence for large-scale internal datacenter systems. The ideal candidate has strong engineering judgment, experience operating complex distributed infrastructure, and the ability to build teams that combine focused operations with automation-first software engineering.

What you will be doing:

  • Lead teams that help us ensure NVIDIA’s internal rack-scale server systems, clusters, and lab facilities remain available, healthy, and reliable.
  • Drive execution across fleet operations, incident response, roadmap planning, change management, operational readiness, and reliability metrics.
  • Build automation, telemetry, alerting, and dashboards that improve visibility and help teams resolve issues faster.
  • Partner with hardware, firmware, software, networking, validation, and infrastructure teams to deploy, sustain, and debug complex systems.
  • Create feedback loops into NPI and sustaining teams to improve product quality, serviceability, and development velocity.
  • Grow and mentor a high-performing technical team with a culture of ownership, learning, and automation-first execution.

What we need to see:

  • BS or MS in Computer Science, Electrical Engineering, Computer Engineering, or related field (or equivalent experience).
  • 12+ overall years of experience in infrastructure, systems engineering, reliability, datacenter operations, distributed systems, or related areas, including 7+ years of people management experience.
  • Strong understanding of server systems, Linux, cluster operations, high-speed networking, and large-scale infrastructure.
  • Experience operating complex systems with high availability expectations, including monitoring, incident management, automation, and fleet-health practices.
  • Proven track record of driving execution across multiple teams, priorities, and technical domains, including close partnership with hardware, firmware, software, networking, validation, and infrastructure organizations.
  • Clear written and verbal communication skills, including executive-level reporting on operational health, risks, and priorities.
  • Track record of building cohesive teams and developing technical leaders who improve reliability and execution.

Ways to stand out from the crowd:

  • Prior Director or Senior Manager experience leading infrastructure, reliability, platform engineering, or large-scale lab operations teams.
  • Experience operating GPU, AI, HPC, cloud, or hyperscale datacenter infrastructure.
  • Broad knowledge of rack-scale systems, including server management, networking, storage, power, thermal, and RAS concepts.
  • Experience building automation, telemetry, fleet health, or dashboarding systems that improve product quality, serviceability, or engineering velocity.

Do you enjoy making complex AI infrastructure reliable at scale while enabling engineering teams to move faster? Come join our datacenter server systems team and help build the reliable, token-efficient computing platforms driving NVIDIA’s success in this exciting and rapidly growing field.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 292,000 USD - 442,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until June 30, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Skills

LinuxChange Management

Similar Jobs

30

Director Operations Engineering

Inspire Medical Systems Inc. · Minneapolis, MN +1

1 week ago

Director, Operations Engineering

Patternenergy · Houston, TX, US

1 week ago

Director Engineering Operations

Astera Labs · Taipei, Taiwan +1

3 weeks ago

Director, Engineering Operations

Abbott · United States > Alameda : 2901 Harbor Bay Parkway, United States of America · Hybrid

3 months ago

Assistant Director, Engineering & Operations

Uasys · UAMS, United States of America · Remote

2 weeks ago

Assistant Director, Engineering & Operations

Career Opportunities · UAMS, United States of America · Remote

2 weeks ago

Sr Director, Operations Engineering

Flextronics · San Jose, California, United States of America

1 month ago

Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems

Nvidia · US, CA, Santa Clara, United States of America

1 month ago

Assistant Director, Engineering & Operations

Uasys · UAMS, United States of America · Remote

1 month ago

Assistant Director, Engineering & Operations

Career Opportunities · UAMS, United States of America · Remote

1 month ago

Solutions Director, Engineering Operations - APAC

JLL employee · SGP-CORP Singapore-PLQ · Onsite

3 months ago

Sr. Director Operations, Engineering

Flextronics · Guadalajara North, Mexico

3 months ago

Director Engineering Operations (Chief of Staff)

Diligentcorporation · New York, New York, United States; Remote U.S; Washington, District of Columbia, United States +3 · Remote

8 months ago

Senior Director, Operations Engineering

Aypapower · United States · Remote

10 months ago

Director, Cybersecurity – Engineering, Operations and Incident Response

Uw Hospital and Clinics · Madison, WI, United States, US · Hybrid

4 days ago

Director of Engineering Operations

High 5 Games

6 days ago

Director, Model Engineering & Operations

CareSource mission is known as · Dayton WFH, United States of America · Remote

1 week ago

Director of Engineering Operations & MES (부장 이상)

JLL employee · KOR-CORP Seoul-JLL Korea HQ, Korea, Republic of · Onsite

1 week ago

Director, Security Engineering & Operations

Payscale · Remote-Canada +1 · Remote

3 weeks ago

Director, Engineering, Platform Operations & Productivity

Gitlab · Remote, Bangalore · Remote

4 weeks ago

Director of Engineering Operations

Spirit Electronics · Phoenix, AZ

1 month ago

Senior Director, IAM Engineering & Operations

UKG · Sunrise, FL,US, US · Remote, Hybrid

2 months ago

Facility Operations and Engineering Director

HealthPartners · Bloomington, MN, United States, US

3 months ago

Director Network Operations Engineering

Depository Trust Company · Jersey City, NJ, United States, US

3 months ago

Director of Engineering & Operations

Action Property Management · San Francisco, CA · Onsite

4 months ago

Director, Cloud Engineering & Operations, FinOps

OneStream Careers Page · Remote, USA · Remote

6 months ago

Director, Utility Operations & Engineering

Employment Opportunities · Boiler Building, United States of America · Onsite

8 months ago

Director of IT Operations & Engineering – MSP

Leapfrog Services Inc · Atlanta, GA

Today

Chief of Transportation Engineering and Construction, Operations Director I (NCS) - Department of Transportation

The City of Baltimore Job Opportunities · Charles L. Benton, Jr. Building, United States of America

3 weeks ago

Director - Advanced Manufacturing Operations Engineering

Generac · Waukesha HQ, United States of America · Onsite

4 weeks ago