Hiring.Camp

Site Reliability Technical Operations Manager

Jci

·

Yesterday

Location
Johnson Controls India COEE1
Type
Full-time
Department
IT
Seniority
Manager
Experience
10+ years
Closing date
Today
Source
Workday

Description

Summary:

The Site Reliability Engineering team at Johnson Controls is seeking a Technical Reliability & Support Manager to lead cloud product reliability, operational support, and production stability across global cloud applications and platforms. This role will be responsible for managing day-to-day technical operations, L2/L3 support coordination, incident response, service reliability governance, and continuous improvement for cloud products hosted across platforms such as Azure (primarily), Google Cloud, Ali Cloud, and other cloud environments.

The role will partner closely with Engineering, SRE, Security, Observability, and external support partners to ensure production issues are resolved quickly, recurring problems are eliminated, support processes are standardized, and operational risks are proactively identified and addressed.


Primary Duties:

  • Lead technical operations and support management for cloud products across global environments
  • Manage day-to-day L2/L3 support activities, incident response, escalations, and production issue resolution
  • Own service reliability governance for assigned cloud products, including availability, incident trends, MTTR, recurring issues, and operational risks
  • Partner with Engineering, SRE, Security, Observability and Platform teams to improve service stability and operational readiness
  • Drive incident management, problem management, RCA/PCA reviews, and corrective action tracking for production issues
  • Ensure timely communication during incidents, including stakeholder updates, executive summaries, customer-impact statements, and resolution updates
  • Establish and manage support processes aligned with ITIL practices, including incident, problem, change, service request, and escalation management
  • Define, track, and report operational KPIs such as availability, MTTR, incident volume, severity trends, backlog, service requests, alert noise, and SLA performance
  • Identify recurring operational pain points and drive permanent fixes through engineering backlog, automation, monitoring improvements, and process enhancements
  • Collaborate with Observability teams to ensure critical application, infrastructure, database, and integration components are properly monitored and alerted
  • Support implementation and adoption of SLIs, SLOs, SLAs, error budgets, and reliability / supportability scorecards for cloud products
  • Ensure production readiness for new product releases, migrations, infrastructure changes, and platform transformations
  • Lead operational reviews with internal teams, vendors, and external support partners to ensure accountability and continuous improvement
  • Drive automation of manual support tasks, ticket workflows, reporting, and operational runbooks to improve efficiency and consistency
  • Ensure support activities, incidents, changes, and action items are properly documented and tracked in tools such as Jira, ServiceNow, Confluence, or equivalent platforms
  • Manage support handoffs across global teams and ensure clear ownership, escalation paths, and communication protocols
  • Provide technical leadership during high-severity incidents, major outages, migrations, and critical customer-impacting events
  • Ensure compliance with security, audit, operational, and IT governance standards for cloud product support
  • Maintain operational documentation, SOPs, runbooks, escalation matrices, support models, and knowledge base articles
  • Mentor support engineers and technical teams on reliability practices, incident handling, RCA quality, and operational excellence

Qualifications:

  • 10+ years of experience in technical operations, production support, SRE, cloud support, or reliability engineering.
  • Strong experience managing cloud-hosted applications and support operations.
  • Good knowledge of Azure (preferred), with exposure to AWS, GCP, or Ali Cloud.
  • Experience with incident management, problem management, change management, and operational governance.
  • Familiarity with cloud-native technologies, microservices, APIs, databases, containers, and Kubernetes.
  • Understanding of observability tools such as Grafana, Datadog, ELK, Logz.io, Azure Monitor, or similar platforms.
  • Knowledge of reliability practices including SLIs, SLOs, SLAs, monitoring, alerting, and automation.
  • Strong troubleshooting skills across applications, infrastructure, networking, databases, and cloud services.
  • Experience with Jira, Confluence, or similar platforms.
  • Excellent leadership, communication, stakeholder management, and vendor coordination skills.
  • Ability to lead high-severity incidents and drive operational excellence in a fast-paced environment.

 Mandatory Skills:

  • Cloud Operations / SRE leadership
  • L2 production support (good to have L2 Production Support)
  • Incident & escalation management
  • RCA/PCA and problem management
  • Azure cloud and cloud-native platforms
  • Observability, monitoring & KPI reporting
  • ITIL-based support processes
  • Automation and runbook development
  • Executive communication & stakeholder management
  • Leadership in high-pressure production environments

Skills

AWSAzureGCPKubernetesJiraConfluenceServiceNowSREMicroservicesComplianceChange ManagementITIL

Similar Jobs

30

Technical Lead - DevOps

Ptc·IND-Pune - Marisoft, India·Remote, Hybrid

1d ago

Technical Lead, Platform Engineering & DevOps

Alcon·Bangalore - AGS, India

2d ago

Technical Lead / Team Lead – DevOps & CI/CD Automation

DXC Technology·AVS01 - DXC Aviles Parque Empresarial, Spain

2d ago

Software Technical Leader - SRE Eventos

MercadoLibre·MG

6d ago

Senior Technical Marketing Engineer, Enterprise Platform

Nvidia·Santa Clara, CA

1w ago

Senior Technical Marketing Engineer, Enterprise Platform

Nvidia·Santa Clara, CA

1w ago

Salesforce DevOps - Technical Architect

Salesforce·India - Bangalore +4·Onsite

1w ago

DevOps Technical Delivery Manager

Barrywehmiller·Raleigh, NC USA +6

2w ago

Technical Software Project Manager – DevOps

Trabajos en Demo Datos·Barcelona

2w ago

Software Technical Leader - Reliability Platforms SRE

MercadoLibre·CABA, AR

2w ago

Senior Staff Site Reliability Operations Technical Lead

Nvidia·Durham, NC·Remote, Hybrid

2w ago

Senior Staff Site Reliability Operations Technical Lead

Nvidia·Durham, NC·Remote, Hybrid

2w ago

Software Engineering Technical Leader - Devops/SRE, AWS, Python/Go - 9 - 14 years

Cisco·IND-BANGALORE, India·Hybrid

2w ago

Senior Technical Account Manager (DevOps)

Perforce Software·Minneapolis, MN·Hybrid

3w ago

DevOps Technical Team Lead, VP

Statestreet·CR1 - 700 District, US

3w ago

Senior Technical Support Engineer (Salesforce Platform)

Model N·Hyderabad India·Hybrid

3w ago

Siftware Engineering Technical Leader - SRE | Security Architect, Kubernetes,AWS,Terraform | 13+ years | Bangalore

Cisco·IND-BANGALORE, India·Remote

1mo ago

Software Engineering Technical Leader - SRE + Kubernetes + Cloud + Automation + Incident Management + AI-first (10-14 Years)

Cisco·IND-BANGALORE, India·Hybrid

1mo ago

Architect / Technical Team Lead - DevOps & DevSecOps

BETSOL·Bengaluru, KA

1mo ago

Technical Systems Engineer – Investment Management Platform, Charles River Development, Assistant Vice President

Statestreet·Dublin 2, Ireland

1mo ago

Senior Technical Consultant - AI Platform Engineer

Ahead·Gurugram, Haryana +1·Remote

1mo ago

Site Reliability Engineering Technical Leader

Cisco·USA-IRVINE, US +6·Remote

1mo ago

Technical Lead Package Platform Engineer

Astera Labs·San Jose, California +1

1mo ago

Sr. Engineer Application and Development and Maintenance - ServiceNow Platform Technical Lead

Cardinalhealth·Philippines-Bonifacio Global City-Taguig

1mo ago

Technical Site Reliability Engineer

Andurilindustries·Abu Dhabi, United Arab Emirates +2

1mo ago

Technical Support Engineer - Platform Support

Talkdesk·Remote·Remote

1mo ago

Technical Lead - Site Reliability Engineering

LSEG·USA-Allen-700 Central Expressway, US +1

1mo ago

Technical Marketing Engineer - AI Platform Software

Nvidia·Santa Clara, CA

2mo ago

Technical Marketing Engineer - AI Platform Software

Nvidia·Santa Clara, CA

2mo ago

Senior System Software Engineer — CPU Platform SBIOS Technical Delivery Lead

Nvidia·Taipei City, TW

2mo ago