Hiring.Camp

Senior Site Reliability Engineer

Electric Power Engineers

·

Yesterday

Salary
Up to $188k
Location
Austin, TX, US
Type
Full-time
Department
IT
Seniority
Senior
Experience
5+ years
Closing date
Today
Source
iCIMS

Description

Overview

We are designing the grid of the future! 

We are seeking an experienced Senior Site Reliability Engineer (SRE) to join our engineering organization. The ideal candidate will combine strong software engineering and cloud infrastructure expertise to improve the reliability, scalability, security, and operational efficiency of our systems. This role will focus on building resilient AWS and Kubernetes platforms, defining and measuring service reliability, automating operational work, and improving how we detect, respond to, and learn from production incidents. 

The Senior SRE will work closely with software engineering, platform, security, QA, and product teams to establish reliability standards and ensure production services can scale safely as the business grows. 

Responsibilities

How you can make an impact:

  • Service Reliability & Availability: Define, measure, and improve service reliability using service-level indicators (SLIs), service-level objectives (SLOs), error budgets, availability targets, and capacity planning.
  • Observability: Build and maintain monitoring, logging, tracing, dashboards, and alerting using tools such as CloudWatch, Prometheus, Grafana, New Relic, or similar platforms. Ensure alerts are actionable and aligned to customer and service impact.
  • Incident Response: Participate in and improve production incident response, including on-call practices, troubleshooting, escalation, communication, root-cause analysis, and blameless post-incident reviews.
  • Operational Excellence: Improve runbooks, documentation, production readiness reviews, change management, operational standards, and engineering practices that reduce risk and improve system maintainability.
  • Cost & Efficiency: Optimize infrastructure for reliability and performance while maintaining responsible cloud spend and supporting FinOps initiatives.
  • Performance & Capacity: Analyze system performance, resource utilization, latency, throughput, and growth trends; identify bottlenecks and implement scalable solutions before they become production issues.
  • Automation & Toil Reduction: Identify repetitive operational work and replace it with reliable automation using Python, Bash, CI/CD tooling, and platform APIs.
  • CI/CD & Release Reliability: Build and improve deployment pipelines that support safe, repeatable releases through automated testing, validation, progressive delivery, rollback strategies, and deployment observability.
  • Security & Compliance: Apply cloud security and operational controls including least-privilege IAM, encryption, network security, secrets management, patching, auditability, and compliance requirements.
  • Infrastructure as Code: Design, develop, review, and maintain reusable Terraform modules and infrastructure-as-code patterns for AWS environments.
  • Kubernetes Reliability: Operate and improve Kubernetes-based platforms, including EKS clusters, workloads, Helm deployments, autoscaling, upgrades, resource management, and workload resilience.
  • Cloud Platform Engineering: Architect, operate, and optimize AWS services such as EC2, S3, RDS, EKS, Lambda, VPC, IAM, Route 53, and related services.
  • Resilience & Disaster Recovery: Design and validate fault-tolerant architectures, backup strategies, recovery procedures, and disaster recovery capabilities. Conduct reliability testing and failure exercises where appropriate.
  • Cross-Functional Collaboration: Partner with development teams to improve application operability, reliability, instrumentation, deployment patterns, and production readiness.

Qualifications

Bring your passion, here's what’s needed:

Required Skills and Qualifications 

  • Experience: 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud infrastructure, or a closely related discipline, including significant production ownership.
  • AWS: Strong hands-on experience designing and operating production workloads in AWS, including networking, IAM, compute, storage, databases, DNS, and managed Kubernetes.
  • Terraform: Advanced proficiency with Terraform, including reusable modules, remote state, dependency management, environment design, code review, and infrastructure lifecycle management.
  • Observability: Experience with metrics, logs, traces, dashboards, alerting, and production telemetry using platforms such as Prometheus, Grafana, CloudWatch, New Relic, ELK/OpenSearch, or similar tools.
  • Reliability Engineering: Practical knowledge of SRE concepts such as SLIs, SLOs, error budgets, capacity planning, fault tolerance, graceful degradation, and reducing operational toil.
  • Kubernetes: Deep understanding of Kubernetes architecture and operations, including EKS, Helm, workload scheduling, networking, storage, autoscaling, upgrades, and troubleshooting.
  • Linux & Systems: Strong Linux systems knowledge and the ability to diagnose issues involving CPU, memory, disk, networking, processes, DNS, and application dependencies.
  • Programming & Automation: Proficiency in Python, Bash, Go, or another general-purpose language used to build operational tooling and automation.
  • Incident Management: Experience troubleshooting complex production incidents and contributing to incident response, root-cause analysis, postmortems, and corrective-action tracking.
  • CI/CD: Experience designing or operating CI/CD systems such as GitHub Actions, Jenkins, GitLab CI, Argo CD, or comparable tooling.
  • Security: Working knowledge of cloud security best practices, including IAM, encryption, secrets management, network segmentation, vulnerability management, and audit controls.
  • Communication: Strong written and verbal communication skills, with the ability to collaborate effectively across engineering and business teams.

Preferred Qualifications 

  • AWS certifications such as AWS Certified Solutions Architect - Professional or AWS Certified DevOps Engineer - Professional.
  • Experience with GitOps practices and tools such as Argo CD or Flux.
  • Experience designing or participating in formal on-call rotations and incident management programs.
  • Familiarity with chaos engineering, resilience testing, or game-day exercises.
  • Experience with service meshes, distributed systems, and microservice architectures.
  • Knowledge of database operations for technologies such as Amazon RDS, DynamoDB, and PostgreSQL.
  • Experience with multi-account AWS environments, landing zones, governance, or large-scale cloud platform design.
  • Experience with infrastructure cost optimization, cloud financial management, or FinOps practices.
  • Familiarity with security and compliance frameworks such as SOC 2, ISO 27001, PCI DSS, or similar standards.

What Success Looks Like 

  • Production services become more measurable, reliable, and resilient over time.
  • Operational toil and recurring incidents are reduced through engineering and automation.
  • Teams have clear SLOs, useful dashboards, actionable alerts, and well-understood operational ownership.
  • Infrastructure and deployment changes are repeatable, observable, secure, and low risk.
  • Incidents produce meaningful learning and durable improvements rather than recurring fixes.
  • Cloud capacity and cost are proactively managed without compromising reliability.

Why Join Us? 

  • Work on modern cloud and reliability engineering challenges in a fast-paced, innovative environment.
  • Help shape reliability standards and engineering practices for mission-critical systems.
  • Collaborate with talented engineers across software, cloud, security, and product disciplines.
  • Competitive salary, comprehensive benefits, and opportunities for professional growth.
  • Flexible remote or hybrid work options.

 

Be a part of an innovative team shaping the grid of the future through advanced energy intelligence.  For more than half a century, Electric Power Engineers (EPE) has partnered with power and energy clients across the globe, providing consulting expertise and energy intelligence software solutions for complex engineering and grid modeling challenges. As leaders in the renewables space, we are focused on building a modern, secure, and resilient grid.   Join us in making an impact on the communities we serve and the environment in which we live. Together we can transform the future of energy.  

 

How we support you:

  • Comprehensive health and wellness benefits including medical, dental, and vision with 100% premium coverage for you
  • Generous PTO and paid holidays
  • MyShare Employee Ownership Program
  • Work with industry leaders
  • 401K, up to a 4% match (100% vested from day 1)

 

 

Location: This position will be located in City, State

Travel:  Occasional travel may be needed (10% or less)

 

EPE is an equal opportunity/AA/Disability/Veteran employer. The EEO is the Law poster, and its supplement are available using the following links: EEOC is the Law Poster

 

 

Third-Party Recruiting Notification

EPE does not accept unsolicited resumes from third-party recruiters. Any unsolicited third-party resumes forwarded by recruiters to EPE via our career page or to any of our managers or employees will be considered public information, may be treated as a direct application from the person identified in the resume, and will not be eligible for placement fee payment to the agency. EPE will not pay a fee to a third-party recruiter or agency without a previously signed third-party agreement and has not coordinated their recruiting activity with the appropriate member of the Talent Acquisition team. 

 

Skills

PythonAWSKubernetesTerraformJenkinsCI/CDLinuxPostgreSQLDynamoDBGitHubGitLabSOCDevOpsSREComplianceChange ManagementSOC 2ISO 27001AWS CertifiedGo

Similar Jobs

30

Senior Platform Engineer

Verisk Careers | Verisk · Malaga, Andalucia, Spain, ES · Hybrid

Today

Senior Platform Engineer

Air New Zealand · Auckland, New Zealand

Today

Senior Platform Engineer

H&M Group · Stockholm, Stockholms län, Sweden

Yesterday

Senior Platform Engineer

9Fin · London

2 days ago

Senior Platform Engineer

Cba · Eveleigh, NSW - 1 Locomotive Street, Australia +1

2 days ago

Senior Platform Engineer

Sherwin-Williams · Cleveland, OH, United States, US

4 days ago

Senior Platform Engineer

DXC Technology · IN323 - Prestige Bangalore EIT - INES,INA7,INET- NSTPI (IN323), India

5 days ago

Senior Platform Engineer

RBA Economic and Finance · Head Office, Australia · Hybrid

5 days ago

Senior Platform Engineer

"HackEDU, Inc. dba Security Journey" · Apex, NC

6 days ago

Senior Platform Engineer

Encompass Corporation · Glasgow · Hybrid

6 days ago

Senior Platform Engineer

Encompass Corporation · Belgrade · Remote

6 days ago

Senior Platform Engineer

Mastercard · London, England (Angel Lane), United Kingdom +2

6 days ago

Senior Platform Engineer

Shippit

6 days ago

Senior Platform Engineer

Ingrammicro · Manila, Philippines

6 days ago

Senior Platform Engineer

Uq · St Lucia Campus, Australia · Hybrid

1 week ago

Senior Platform Engineer

UQ Careers · St Lucia Campus, Australia · Hybrid

1 week ago

Senior Platform Engineer

We are hiring! · Warsaw, Poland · Remote, Hybrid

1 week ago

Senior Platform Engineer

Focused · Denver, Colorado, United States +1

1 week ago

Senior Platform Engineer

Trust Bank · Singapore +1

1 week ago

Senior Platform Engineer

LSEG · St. Loui, Missouri, United States of America

1 week ago

Senior Platform Engineer

DXC Technology · MX446 - DXC Mexico City Lago Alberto Wework (MX446)

1 week ago

Senior Platform Engineer

Webai · Austin, TX

1 week ago

Senior Platform Engineer

Goswift · Europe · Remote

1 week ago

Senior Platform Engineer

Clarity Innovations · Augusta, GA; San Antonio, TX; Ft Meade, MD +3 · Onsite

1 week ago

Senior Platform Engineer

Impact · Cape Town +1

1 week ago

Senior DevOps

Intive · Remote: Argentina · Remote

2 weeks ago

Senior Platform Engineer

Kojo · Mexico · Remote

2 weeks ago

Senior Platform Engineer

Thomson Reuters · United States of America, Eagan, Minnesota +3 · Hybrid

2 weeks ago

Senior Platform Engineer

LSEG · St. Loui, Missouri, United States of America

2 weeks ago

Senior Platform Engineer

Scale Army Careers · Egypt +4 · Remote

2 weeks ago
Senior Site Reliability Engineer at Electric Power Engineers • Up to $188k | Hiring.Camp