Hiring.Camp

Senior Manager, Reliability Engineering & AIOps

Lam Research

·

Today

Location
Bengaluru, KA,IN, IN
Type
Full-time
Department
Engineering
Seniority
Senior
Experience
12+ years
Education
Master
Source
Eightfold

Description

Lead, hire, and develop the reliability engineering team, owning on-call health while staying technically hands-on. Set the reliability strategy: define the service level objective program, publish an error-budget policy, and drive adoption across platform and service teams. Build and run a follow-the-sun on-call and response model across six regions, with clean handoffs and one consistent set of runbooks and severity definitions worldwide Own the incident management and paging platform end to end, including services, schedules, escalation policies, and routing, configured as code and tuned so alerts fire on real risk rather than noise Serve as incident commander on major incidents, own executive and stakeholder communications, and lead blameless postmortems with tracked follow-up. Own disaster recovery strategy and execution across Azure, AWS, GCP, and core infrastructure platforms, including service-tier recovery objectives, backup and restore validation, failover readiness, DR certification, runbook governance, and recurring exercises measured against RTO and RPO targets Lead capacity planning and performance engineering across Azure, AWS, GCP, compute, storage, network, and HPC platforms, using demand forecasting, utilization trends, growth modeling, and automation to prevent capacity risk and reduce manual operational work Define and drive AI Ops requirements for reliability engineering across Azure, AWS, and GCP, including Microsoft Copilot, Cursor, GitHub Copilot, and LLM-based operational workflows for incident triage, runbook generation, knowledge retrieval, root-cause analysis, and safe remediation recommendations. This is a full-time role on a standard schedule, with participation in a global on-call rotation. Bachelor's degree in Computer Science, Engineering, or a related field with 15 years of related experience; or a Master's degree with 12 years of experience; or equivalent experience. Experience leading or mentoring a reliability or operations team and setting technical direction Proven incident command on major outages, plus ownership of a postmortem process Strong background in disaster recovery planning across Azure, AWS, GCP, and core infrastructure platforms, including restore validation, failover testing, recovery-objective definition, and corrective action tracking after DR exercises or production incidents Hands-on ownership of an incident management and paging platform at scale, such as PagerDuty Experience with capacity planning, performance trending, utilization analysis, and infrastructure demand forecasting for globally distributed production environments across Azure, AWS, GCP, and on-premises platforms Track record of defining and defending service level objectives and error budgets in production Working depth in observability tooling (Prometheus, Grafana, Loki, Tempo or equivalent), infrastructure as code (Terraform), and Python or Go Practical experience applying AI-assisted engineering and operations tools such as Microsoft Copilot, Cursor, GitHub Copilot, or enterprise LLM platforms to improve troubleshooting, automation, documentation, and engineering productivity across Azure, AWS, GCP, and hybrid infrastructure, with clear guardrails for security, privacy, auditability, and production safety. Experience running global, follow-the-sun operations across multiple regions and time zones. Capacity and performance engineering at multi-region scale, including Azure, AWS, GCP, high-performance computing, large storage estates, hybrid cloud infrastructure, and proactive capacity governance. Policy as code, progressive delivery, and chaos engineering in practice. Experience building or operating AI and agent-assisted automation in operations, with a clear view of its failure modes. Experience designing or operating AI Ops capabilities across Azure, AWS, GCP, and hybrid environments, including LLM-grounded knowledge bases, agent-assisted incident workflows, prompt and evaluation practices, and supervised automation that can recommend or propose operational changes before execution.

Skills

PythonAWSAzureGCPTerraformGitHub

Similar Jobs

30

Senior Manager- Reliability

First Solar · Perrysburg, OH, United States, US

4 days ago

Senior Manager, Reliability Engineering & AIOps

Lam Research · Fremont, CA,US, US

Yesterday

Senior Manager, Site Reliability Engineering (SRE)

Nium · Bangalore +1 · Hybrid

2 days ago

(USA) Senior Manager I, Reliability

Walmart · (USA) IN MCCORDSVILLE 08972 ECOMM FULFILLMENT SERVICES, United States of America

3 days ago

Senior Manager, Reliability (GERA)

Dayone · CFA Campus (SG), Singapore

3 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · China · Remote

5 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · Indonesia · Remote

5 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · United Kingdom · Hybrid

5 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · Seoul, Korea · Hybrid

5 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · Singapore · Remote

5 days ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

Bjakcareer · United States · Remote

5 days ago

Senior Manager, Reliability & Platform Engineering

McMaster-Carr · Chicago, IL (Elmhurst) +1 · Hybrid

1 week ago

Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)

84.51° · Cincinnati, OH +2

1 week ago

Senior Manager, Site Reliability Engineering

Oracle · Ireland, IE

1 week ago

Senior Manager, Service Reliability and Customer Operations

InterSystems · Dublin

1 week ago

Senior Manager, Site Reliability Engineering – Paylo Platform

PDI Technologies · Alpharetta, GA +3 · Hybrid

2 weeks ago

Senior Manager, Site Reliability & Operational Resilience

Zelis Careers · US NJ Morristown, United States of America +4

2 weeks ago

Senior Manager, Maintenance & Reliability Engineering

Avantor · USA-NJ Phillipsburg, United States of America

2 weeks ago

Senior Manager, Site Reliability Engineering

Oracle · Nashville, TN, United States, US

2 weeks ago

Senior Manager, Site Reliability & Infrastructure Engineering

Aviva the most attractive choice · Canada - Markham ON 10 Aviva Way · Hybrid

3 weeks ago

Senior Manager, Data & AI Platforms Reliability

AmeriLife has served the needs · Remote, FL, United States of America +47 · Remote

3 weeks ago

Senior Lab Manager - Reliability

Vertiv · Pune, India

3 weeks ago

Senior Product Manager, Reliability & Call Quality

Tin Can · Seattle · Hybrid

3 weeks ago

Senior Manager System Reliability Engineering

Gevernova · Hyderabad TS IN 26, India · Hybrid

3 weeks ago

Senior Manager System Reliability Engineering

gevernova · Hyderabad TS IN 26, India · Hybrid

3 weeks ago

Senior Manager, Hardware Reliability & Test

Muonspace · San Jose, CA +1 · Hybrid, Onsite

4 weeks ago

Senior Manager, Site Reliability Engineering

Finastra · Mississauga - Avebury, Canada

1 month ago

Senior Engineering Manager - Enterprise Trust & Reliability

Multiverse · London · Hybrid

1 month ago

Senior Maintenance and Reliability Manager

Abbott · United States > Casa Grande : Plant, United States of America

1 month ago

Senior Maintenance and Reliability Manager

Abbott · United States > Casa Grande : Plant, United States of America

1 month ago