Hiring.Camp

Senior Staff Site Reliability Engineer – Compute Platform

Nvidia

·

Today

Location
India, Bengaluru
Type
Full-time
Department
Engineering
Seniority
Senior
Source
Workday

Description

NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services.

What you’ll be doing:

  • Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale.

  • Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation.

  • Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data.

  • Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems.

  • Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation.

What we need to see:

  • BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services.

  • Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges.

  • Experience deploying and operating bare-metal infrastructure in a data-center environment, including provisioning, networking, operating-system lifecycle management, and hardware automation.

  • Proficiency in Python, Go, or a comparable programming language, with experience building RESTful services and integrating infrastructure APIs.

  • Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, Chef, or Puppet, along with a solid understanding of TCP/IP networking and infrastructure security.

  • Strong SRE and observability experience, including SLIs, SLOs, error budgets, incident management, monitoring, logging, tracing, and tools such as OpenTelemetry, Prometheus, Grafana, ELK Stack, or Splunk.

  • Clear written and interpersonal communication skills, with a record of delivering practical, scalable solutions to complex technical problems.

Ways to stand out from the crowd:

  • Experience operating HPC, AI, GPU-accelerated, or general-purpose bare-metal compute infrastructure, including GPU-enabled Kubernetes or KubeVirt clusters.

  • Expertise with VMware vSphere, Red Hat OpenShift, KVM, Firecracker, OpenStack, or Nutanix AHV.

  • Experience applying generative AI or agentic workflows to improve infrastructure diagnostics, reduce operational toil, and accelerate incident resolution.

  • Experience building secure, integrated operational platforms using APIs, RBAC, service accounts, secrets management, audit controls, workflow orchestration, and infrastructure or incident-management systems.

  • Demonstrated delivery of complex, high-impact infrastructure projects.

Skills

PythonDockerKubernetesTerraformAnsibleLinuxVMwareSplunkTCP/IPSREMicroservicesGo

Similar Jobs

30

Sr. Staff Site Reliability Engineer

Zscaler · Bangalore, IND · Remote, Hybrid

Today

Senior Staff Site Reliability Engineer – Compute Platform

Nvidia · Bengaluru, KA,IN, IN

Today

Senior Staff Engineer - Civil/Site

French & Parrello Associates · Wall Township, NJ

5 days ago

Senior Staff Site Reliability Operations Technical Lead

Nvidia · US, NC, Durham, United States of America · Remote, Hybrid

6 days ago

Senior Staff Site Reliability Operations

Nvidia · US, WA, Seattle, United States of America · Remote, Hybrid

6 days ago

Senior Staff Site Reliability Operations

Nvidia · Seattle, WA,US, US · Remote, Hybrid

6 days ago

Senior Staff Site Reliability Operations Technical Lead

Nvidia · Durham, NC,US, US · Remote, Hybrid

6 days ago

Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)

Gitlab · Remote · Remote

1 week ago

Senior Staff Engineer - Site Reliability

Freshworks CRM · Hyderabad, TS, India

1 week ago

Senior Staff Engineer - Site Reliability

Freshworks CRM · Chennai, India

1 week ago

Senior Staff Site Reliability Engineer

Nvidia · India, Bengaluru +1

1 week ago

Senior Staff Site Reliability Engineer

Nvidia · Bengaluru, KA,IN, IN +1

1 week ago

Sr. Staff Lead Site Reliability Engineer (R5803)

Shieldai · San Diego, California · Onsite

1 week ago

Sr. Staff Lead Site Reliability Engineer (R5803)

Shieldai · San Mateo, California · Onsite

1 week ago

Senior Staff Site Reliability Engineer

Nvidia · India, Bengaluru

2 weeks ago

Senior Staff Site Reliability Engineer

Nvidia · Bengaluru, KA,IN, IN

2 weeks ago

Senior Staff Site Reliability Engineer

Servicetitan · India Bengaluru, Karnataka

3 weeks ago

Senior Staff Site Reliability Engineer

Nvidia · Bengaluru, KA,IN, IN

4 weeks ago

Senior Staff Site Reliability Engineer

Nvidia · India, Bengaluru

4 weeks ago

Senior Staff Site Reliability Engineer

Ping Identity · USA - Remote +1 · Remote

1 month ago

Senior Staff Site Reliability Engineer

Nvidia · Bengaluru, KA,IN, IN

1 month ago

Senior Staff Site Reliability Engineer

Nvidia · India, Bengaluru

1 month ago

Sr/Staff Site Reliability Engineer, Consumer Apps

Attain · Chicago, IL +1 · Remote, Hybrid, Onsite

2 months ago

Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms

Gitlab · Remote · Remote

2 months ago

Sr. Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)

Zscaler · Hyderabad, IND

2 months ago

Sr Staff Site Reliability Engineer (SRE)

Arrow · Ahmedabad, India +4

3 months ago

Senior Staff Site Reliability Engineer

Ironcladhq · San Francisco +2 · Hybrid

3 months ago

Senior Staff Site Reliability Engineer

Hivewatch · El Segundo, CA +1

3 months ago

Senior Staff Site Reliability Engineer

LiveRamp is the data collaboration · San Francisco, United States of America

4 months ago

Sr. Staff Site Reliability Engineer-Federal, Security Clearance

Zscaler · Crystal City, Virginia, USA +1 · Remote

4 months ago