Hiring.Camp

Staff Engineer, Site Reliability Engineering

General Motors

·

Yesterday

Location
GM Global Technical Center - Cole Engineering Center Tower, United States of America
Workplace
Hybrid
Type
Full-time
Department
Engineering
Seniority
Senior
Education
PhD
Source
Workday

Description

Job Description

About the role


General Motors is transforming the automotive landscape through its next-generation Software-Defined Vehicle platform. Data is central to that transformation, powering safety, personalization, energy optimization, operational decision-making, and connected customer experiences.

We are seeking a Staff Engineer to help make GM’s data platforms reliable, observable, operable, and scalable. This is a senior technical leadership role for someone who can move comfortably between system-level design, production operations, incident response, automation, and customer partnership.

You will help define and spread the engineering patterns that make services easier to operate. You will work with SRE, data engineering, infrastructure, developer experience, application, and product teams to improve reliability from design through production and continuously improve how the organization operates.


What you’ll do

  • Lead the design and implementation of scalable, fault-tolerant, and observable infrastructure supporting vehicle telemetry, data ingestion, and platform operations.
  • Lead production readiness efforts across multiple teams—engaging directly in code, shaping reliability standards, guiding architectural improvements, and ensuring applications launch with resilient deployments, strong observability, and predictable operations.
  • Design, implement, and improve CI/CD delivery pipelines that make releases repeatable, safe, observable, and fast. Establish appropriate quality gates, artifact promotion, deployment verification, progressive delivery, and rollback practices.
  • Partner across SRE, product, and application teams to design and implement meaningful SLOs, SLIs, observability, monitoring and alerting, runbooks, and operational best practices.
  • Build and improve reusable AI workflows, skills, and evaluations. Apply appropriate validation techniques, including regression testing, structured evaluations, and LLM-as-a-judge approaches where useful.
  • Automate operational work, including incident intake, triage, diagnostics, remediation, evidence collection, service requests, and customer-facing status workflows.
  • Participate in a weekly on-call rotation with 12-hour shifts; the rotation cycles every eight weeks. Lead incident response, communicate clearly under pressure, and coordinate effective mitigation and recovery.
  • Participate in post-incident reviews and drive durable, system-level fixes that prevent recurrence rather than relying on short-term patches or repeated manual workarounds.
  • Partner directly with internal customers to understand their needs, explain technical trade-offs, and improve service outcomes with tact, empathy, and clear communication.
  • Influence technical direction across teams, mentor engineers, and raise engineering standards through design reviews, code reviews, documentation, and hands-on leadership with cross-functional engineering projects.
  • Balance reliability, performance, security, delivery speed, and cost when making technical decisions—especially under pressure.

What you bring

  • 8+ years in SRE, DevOps, or systems engineering, including experience managing or mentoring high-impact teams.
  • Track record of designing and building and maintaining high-scale, cloud-native systems in production (preferably Azure, AWS, or GCP).
  • Hands-on experience architecting observability patterns, including standardized instrumentation, OTEL collector configuration, SLO/SLI definitions, and deploying observability resources like monitors, alerts, and dashboards.
  • Strong understanding of production readiness, service ownership, SLOs, incident management, post-incident learning, and continuous reliability improvement.
  • Experience participating in an on-call rotation and leading technical response to production incidents.
  • Experience designing, operating, and improving CI/CD pipelines. Understanding of GitOps, release strategies, quality gates, deployment verification, progressive delivery, and safe rollback is expected.
  • Strong programming ability in Python, Go, Java, or a comparable language, with disciplined code review, version control, testing, and maintainability practices.
  • Familiarity with AI-assisted software development, LLM application practices, agentic workflows, and evaluation techniques.
  • Ability to work effectively with internal customers, including in difficult or high-pressure situations, with professionalism, tact, and empathy.
  • Excellent ownership attitude and the ability to operate with pace, judgment, and accountability in a high-velocity environment.
  • Ability to influence without relying on formal authority and to increase adoption of shared engineering patterns.
  • Strong written and verbal communication skills for both technical and non-technical audiences.
  • BS / MS / PhD in computer science, engineering, physics, mathematics, or another relevant, technical field

Preferred experience

  • Azure Databricks
  • Azure Event Hubs
  • Azure Kubernetes Service (AKS)
  • Kubernetes configuration management with Helm and Kustomize
  • Infrastructure as code, especially Terraform
  • GitHub Actions, Argo CD, and GitOps-based deployment models
  • Prometheus, Grafana, Datadog, OpenTelemetry, or comparable observability platforms and tools
  • LLM application development and testing with Promptfoo, agentic workflows and reusable AI skills with CoPilot
  • Experience operating large-scale data ingestion, processing, and delivery systems such as Fivetran, Apache Flink, Kafka, and Pulsar
  • Experience with vehicle telemetry, connected-vehicle platforms, or other high-volume event-driven systems

Why join us?

This is an opportunity to shape the reliability foundations of GM’s next generation of connected, software-defined vehicles. You will work on meaningful systems at global scale while helping define how modern SRE teams use automation, observability, operational excellence, and AI to deliver better outcomes for customers.

GM does not provide immigration-related sponsorship for this role. Do not apply for this role if you will need GM immigration sponsorship now or in the future. This includes direct company sponsorship, entry of GM as the immigration employer of record on a government form, and any work authorization requiring a written submission or other immigration support from the company (e.g., H1-B, OPT, STEM OPT, CPT, TN, J-1, etc). This role is categorized as hybrid. This means the selected candidate is expected to report to a specific location at least 3 times a week {or other frequency dictated by their manager}. This job is not eligible for relocation benefits. Any relocation costs would be the responsibility of the selected candidate.

About GM

Our vision is a world with Zero Crashes, Zero Emissions and Zero Congestion and we embrace the responsibility to lead the change that will make our world better, safer and more equitable for all.

Why Join Us 

We believe we all must make a choice every day – individually and collectively – to drive meaningful change through our words, our deeds and our culture. Every day, we want every employee to feel they belong to one General Motors team.

Benefits Overview

From day one, we're looking out for your well-being–at work and at home–so you can focus on realizing your ambitions. Learn how GM supports a rewarding career that rewards you personally by visiting Total Rewards resources.

Non-Discrimination and Equal Employment Opportunities (U.S.)

General Motors is committed to being a workplace that is not only free of unlawful discrimination, but one that genuinely fosters inclusion and belonging. We strongly believe that providing an inclusive workplace creates an environment in which our employees can thrive and develop better products for our customers.

All employment decisions are made on a non-discriminatory basis without regard to sex, race, color, national origin, citizenship status, religion, age, disability, pregnancy or maternity status, sexual orientation, gender identity, status as a veteran or protected veteran, or any other similarly protected status in accordance with federal, state and local laws. 

We encourage interested candidates to review the key responsibilities and qualifications for each role and apply for any positions that match their skills and capabilities. Applicants in the recruitment process may be required, where applicable, to successfully complete a role-related assessment(s) and/or a pre-employment screening prior to beginning employment. To learn more, visit How we Hire.

Accommodations

General Motors offers opportunities to all job seekers including individuals with disabilities. If you need a reasonable accommodation to assist with your job search or application for employment, email us or call us at 1-800-865-7580. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

Skills

PythonJavaAWSAzureGCPKubernetesTerraformCI/CDDatabricksData EngineeringGitHubDevOpsSREGo

Similar Jobs

30

Staff Site Reliability Engineer, Payments Infrastructure

Hireworksio·Brasília

3d ago

Sr. Staff Site Reliability Engineer-Federal, Security Clearance

Zscaler·Crystal City, Virginia +1·Remote, Onsite

3d ago

Staff Site Reliability Engineer (Remote)

Knowbe4·Remote·Remote

1w ago

Staff Site Reliability Engineer (Remote in Brazil)

Knowbe4·São Paulo, Brazil

1w ago

Senior/Staff Site Reliability Engineer

Factorial·A Coruña, ES·Onsite

1w ago

Staff Site Reliability Engineer, Tech Lead

Guidewire·Bengaluru Office, India·Hybrid

1w ago

Sr Staff Site Reliability Engineer, AI Infrastructure

D Matrix·Santa Clara·Hybrid

1w ago

Staff Site Reliability Engineer (f/m/d)

1&1 Drillisch·Revaler Straße 28-31, 10245 Berlin +1·Remote, Hybrid

1w ago

Senior/Staff Site Reliability Engineer

Factorial·Barcelona, ES·Onsite

2w ago

Staff Site Reliability Engineer

Horizon3Ai·Remote, US·Remote

2w ago

Staff Site Reliability Engineer

Andurilindustries·Costa Mesa, California

2w ago

Staff Site Civil Engineer

The Kleingers Group·West Chester, OH +1

2w ago

Staff Manufacturing Engineer (Grecia, Manufacturing Site)

Stryker is one of the·Alajuela, Grecia Evolution Free Zone Building 2·Onsite

2w ago

Senior Staff Site Reliability Engineer

Nvidia·India, Bengaluru

2w ago

Senior Staff Site Reliability Engineer

Nvidia·Bengaluru, KA

2w ago

Staff Site Reliability Engineer - Federal

ServiceNow·San Diego, CALIFORNIA·Hybrid

2w ago

Staff Site Reliability Engineer

Renesas Electronics·La Jolla, CALIFORNIA·Hybrid

2w ago

Staff Site Reliability Engineer

Danaher·POL – Krakow – Cytiva, Poland·Remote, Onsite

2w ago

Staff Site Reliability Engineer

Replit·Remote - United States·Remote

3w ago

Sr. Staff Site Reliability Engineer

Zscaler·Bangalore, IND·Remote, Hybrid

3w ago

Staff Engineer, Continuous Improvement (On-site)

Stryker is one of the·Puerto Rico, Arroyo Las Guasimas Industrial Park·Onsite

3w ago

Staff Databricks Engineer, Site Reliability Engineering

Kinaxis·Perungudi, TN

4w ago

Staff Site Reliability Engineer

Earnin·Mountain View, US +1·Hybrid, Onsite

4w ago

Staff Site Reliability Engineer

ServiceNow·Dublin, Ireland·Hybrid

4w ago

Staff Site Reliability Engineer

Crunchyroll·Los Angeles, California +1

1mo ago

Staff Site Reliability Engineer

Crunchyroll·San Francisco, CA +1

1mo ago

Entry Level Staff Engineer - Site Civil Design - Summer 2027

Tighebond·Sandwich, MA·Remote

1mo ago

Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)

Gitlab·Remote·Remote

1mo ago

Staff Site Reliability Engineer - AI Platform Runtime

Nvidia·Santa Clara, CA

1mo ago

Staff Site Reliability Engineer - AI Platform Runtime

Nvidia·Santa Clara, CA

1mo ago