Hiring.Camp

Sr. Site Reliability Engineer

Octanner

·

1 week ago

Location
USA - Utah-Salt Lake City-Headquarters, United States of America
Type
Full-time
Department
Engineering
Experience
5+ years
Source
Workday

Description

O.C. Tanner is the global leader in software and services that improve workplace culture through meaningful employee experiences. Our Culture Cloud is a suite of apps designed to enhance the employee experience with strategic recognition, service awards, wellbeing, leadership, and events that help people thrive at work. Our Culture by Design approach provides expert services to organizations looking to create great workplaces.

Our global team of 1,500 people hail from 58 countries and speak 62 languages. As programmers, researchers, designers, client professionals and craftspeople we create the tech, tools and awards that connect employees to purpose at thousands of companies. Join us as we help people all over the world thrive at work.

Location: Salt Lake City, UT

As a Senior Site Reliability Engineer, you will help define the future of reliability for our world-class employee recognition platform. You'll leverage software engineering, automation, and cloud-native technologies to build and operate highly available, scalable systems that serve millions of users. We're looking for someone who is passionate about reliability engineering, continuous improvement, and building self-healing platforms that enable development teams to move faster while delivering exceptional customer experiences.

Key Responsibilities

  • Improve the availability, scalability, and performance of cloud-native applications through automation, monitoring, and engineering best practices.
  • Build and evolve observability platforms using OpenTelemetry, Datadog, Coralogix, or similar tools. Establish standards for metrics, logs, traces, and service-level objectives (SLOs) that enable proactive issue detection and resolution.
  • Lead production triage efforts, rapidly diagnosing and resolving service disruptions. Drive incident management, root cause analysis, and blameless post incident reviews to improve system resilience and reduce recurring issues.
  • Partner with Engineering, Support and Product teams to embed reliability, observability, and operational excellence throughout the software development lifecycle.
  • Champion a reliability-first engineering culture by establishing automation standards, monitoring best practices, shift-left quality approaches and shared ownership models that proactive improve resilience, reduce operational risk, and protect the availability of business-critical services.
  • Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective handoffs, and shared ownership of production services.
  • Participate in an on-call rotation focused on maintaining service health, reducing operational toil, improving alert quality, and automating repetitive operational tasks.

 Required Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, or related roles, with a strong background in production triage, incident response, and operational excellence.
  • Experience operating large-scale, customer-facing SaaS platforms with high availability and uptime requirements.
  • Proficiency in Go, Python, Java, or similar programming languages, with demonstrated experience building automation, production tooling, and reliability-focused engineering solutions.
  • Deep experience with modern Infrastructure-as-Code and GitOps technologies such as Terraform, OpenTofu, CDKTF, Pulumi, ArgoCD, Helm, and Kubernetes.
  • Hands-on experience with OpenTelemetry, Datadog, Coralogix, or similar observability platforms.
  • Strong knowledge of AWS services and Kubernetes in production environments.
  • Deep understanding of monitoring, logging, and distributed tracing for complex systems.
  • Ability to partner effectively with software engineering and testing teams to design reliable systems, improve application performance, and strengthen quality practices across the software development lifecycle.
  • Comfortable with participating in on-call rotations and handling high-pressure environments.

Bonus Qualifications:

  • Experience with multiple cloud or cloud-agnostic environments.
  • Familiarity with security, compliance, and governance frameworks
  • Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
  • Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.

Skills

PythonJavaAWSKubernetesTerraformPostgreSQLRedisDevOpsCompliance

Similar Jobs

30

Sr. Site Reliability Engineer

Illumio · HQ - Sunnyvale (Office) · Onsite

3 days ago

Sr. Site Reliability Engineer

Illumio · HQ - Sunnyvale (Office) · Onsite

1 week ago

Sr. Site Reliability Engineering (Agentic Builders Experience team)

Adobe · San Jose, United States of America

1 week ago

Sr. Site Reliability Engineer

Versant · Orlando, FL, United States · Remote

2 weeks ago

Sr. Site Reliability Engineer

Fiserv is the global leader · Berkeley Heights, New Jersey, United States of America +2 · Onsite

2 weeks ago

Sr Site Reliability Engineer

Workday · USA, GA, Atlanta, United States of America

2 weeks ago

Sr. Site Reliability Engineer - Paze

earlywarningservices · Scottsdale, United States of America +2 · Hybrid

3 weeks ago

Sr. Site Reliability Engineer (Application Software)

Spacex · Hawthorne, CA +1

3 weeks ago

Sr Site Reliability Engineer - Release

Alkami Careers · Home Office, United States of America · Remote

3 weeks ago

Sr Site Reliability Engineer I

Disney · IND - Bangalore - Ecoworld 6B - 7th Floor, India · Onsite

3 weeks ago

Sr. Site Reliability Engineer, Data - FreeWheel

Comcast · IL - Chicago, 350 N. Orleans St 1300N, United States of America

3 weeks ago

Sr. Site Reliability Engineer, Data - FreeWheel

Comcast · VA - Reston, 11951 Freedom Dr Ste 900, United States of America

3 weeks ago

Sr Site Reliability Engineer

Rdccareers · Austin, Texas, United States

3 weeks ago

Sr. Site Reliability Engineer - SRE

" QAD, Inc." · Barcelona, CT, Spain · remote

3 weeks ago

Sr Site Reliability Engineer

Signoz · India · Remote

3 weeks ago

Sr. Site Reliability Engineer

Spacex · Washington, DC +2

1 month ago

Sr. Site Reliability Engineer

Versant · Orlando, FL, United States · Remote

1 month ago

Sr. Site Reliability Engineer - SRE

" QAD, Inc." · Barcelona, CT, Spain · Remote

1 month ago

Sr. Site Reliability Engineer (Starshield)

Spacex · Washington, DC +3

1 month ago

Sr. Site Reliability Engineer (Starshield)

Spacex · Redmond, WA +3

1 month ago

Sr. Site Reliability Engineer (Starshield)

Spacex · Hawthorne, CA +3

1 month ago

Digital Site Reliability Sr Engineer - Remote

NTT DATA · Memphis, TN,US, US

1 month ago

Sr Site Reliability Engineer

Renaissancelearning Nam · Remote- US +1 · Remote

1 month ago

Sr. Site Reliability Engineer

Mike Albert Fleet Solutions · Cincinnati, OH

1 month ago

(Sr) Site Reliability Engineer (US Federal)

Workday · USA.VA.Reston, United States of America · Onsite

1 month ago

Sr Site Reliability Engineer

Adobe · Bucharest, Romania

1 month ago

Sr. Site Reliability Engineer

Illumio · Sunnyvale, California - HQ · Onsite

1 month ago

Sr. Site Reliability Engineer

Illumio · Sunnyvale, California - HQ · Onsite

1 month ago

Sr. Site Reliability Engineer

Pitchbookdata · Seattle, Washington, United States

1 month ago

Sr. Site Reliability Engineer

Pitchbookdata · Seattle, Washington, United States

1 month ago