Hiring.Camp

Staff Site Reliability Engineer- Eng

UKG

·

Jun 22, 2026

Location
Lowell, MA,US, US
Type
Full-time
Department
Engineering
Seniority
Senior
Experience
5+ years
Source
Eightfold

Description

Engage in and improve the lifecycle of services from conception to end-of-life, including system design reviews, capacity planning, and production readiness. Define and implement standards and best practices for system architecture, service delivery, reliability, and automation, including the definition and monitoring of service health indicators (latency, traffic, error rates, and resource saturation), service level objectives (SLOs), and the use of error budgets to guide operational and delivery decisions. Support service, product, and engineering teams by providing common tooling and frameworks to increase availability and improve incident detection and response. Improve system performance, availability, and efficiency through automation, process refinement, post-incident reviews, and in-depth configuration analysis. Collaborate closely with engineering teams across the organization to deliver and operate reliable services. Increase operational efficiency, effectiveness, and service quality by treating operational challenges as software engineering problems (reducing toil). Guide junior team members and serve as a champion for Site Reliability Engineering best practices. Actively participate in incident responses, including on-call rotations and post-incident reviews, collaborating with engineering teams to restore service and reduce recurrence. Partner with stakeholders to influence and help drive the best possible technical and business outcomes. 5+ years of hands-on experience in software engineering, systems engineering, or cloud-based environments. 5+ years of experience working with public cloud platforms (e.g., GCP (preferred), AWS, or Azure). 5+ years of experience configuring, operating, and maintaining applications and/or systems infrastructure in a large-scale, customer-facing environment. Demonstrated understanding of observability best practices, including metric generation and collection, log aggregation pipelines, time-series databases, and distributed tracing. Experience coding in one or more higher-level programming languages (e.g., Python, Java, or C++). Strong working knowledge of Linux systems, including troubleshooting, performance analysis, and scripting in production environments. Experience with GitHub Actions and modern CI/CD practices. Experience building operational dashboards and alerts using observability tools such as Splunk or Grafana. Excellent communication and collaboration skills, with experience of mentoring and guiding engineers. Hands-on experience with cloud-native applications and containerization technologies (Kubernetes, containers). Experience with infrastructure-as-code and configuration management tools (e.g., Terraform, Ansible). Experience operating production workloads in Google Cloud Platform (GCP). Solid grounding in at least two of the following areas: Computer Science fundamentals, Cloud Architecture, Security, or Network Design.

Skills

PythonJavaAWSAzureGCPKubernetesTerraformAnsibleCI/CDLinuxGitHubSplunk

Similar Jobs

30

Staff Site Reliability Engineer (SRE) (Hybrid)

Cisco · USA-SAN FRANCISCO, United States of America +10 · Hybrid

Today

Staff Site Reliability Engineer

Genomics · London +1 · Hybrid

Yesterday

Staff Site Reliability Engineer

360 Privacy · Brentwood, Tennessee

Yesterday

Staff Site Reliability Engineer

BeyondTrust · Remote Canada | Remote United States +1 · Remote

2 days ago

Staff Site Reliability Engineer

Circle · U.S. - California, United States of America · Remote

3 days ago

Staff Site Reliability Engineer

Gradle Technologies · Europe (GMT) +1

1 week ago

Staff Site Reliability Engineer- Eng

UKG · Noida, UP,IN, IN

1 week ago

Staff Site Reliability Engineer

Trimble · Chennai, TN,IN, IN

1 week ago

Staff Site Reliability Engineer

Trimble · India - Chennai

1 week ago

Staff Site Reliability Engineer (AI Platform)

Manychat · Barcelona, Spain +1 · Hybrid

1 week ago

Staff Site Reliability Engineer (AI Platform)

Manychat · Amsterdam, Netherlands +1

1 week ago

Staff Site Reliability Engineer

Guidewire · Kuala Lumpur Office, Malaysia · Hybrid

2 weeks ago

Staff Site Reliability Engineer

Pingidentity · UK - Remote +1 · Remote

2 weeks ago

Staff Site Reliability Engineer

Ionq · Santa Clara, California, United States

3 weeks ago

Sr/Staff Site Reliability Engineer, Consumer Apps

Attain · Chicago, IL +1 · Remote, Hybrid, Onsite

3 weeks ago

Staff Site Reliability Engineer

Servicetitan · India Bengaluru, Karnataka

3 weeks ago

Staff Site Reliability Engineer

Tenex · Remote, USA · Remote

3 weeks ago

Sr. Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)

Zscaler · Hyderabad, IND

1 month ago

Staff Site Reliability Engineer

Stryker is one of the · Haryana, Gurugram International Techpark, Block I Phase 1 Floors G, 3, 4, 5, India · Hybrid

1 month ago

Staff Site Reliability Engineer

Chamberlain · Oak Brook, United States of America

1 month ago

Staff Site Reliability Engineer

Stackblitz · Remote · Remote

1 month ago

Staff Platform Engineer / Staff Site Reliability Engineer

Andurilindustries · Sydney, New South Wales, Australia

1 month ago

Staff Site Reliability Engineer - Paze

earlywarningservices · San Francisco, United States of America +2 · Hybrid

1 month ago

Staff Site Reliability Engineer

The Onset Jobs Marketplace · AU

1 month ago

Staff Site Reliability Engineer - Volcano

Kong · United States · Remote

1 month ago

Staff Site Reliability Operations Engineer

Calix will come from · Remote - USA, United States of America · Remote

1 month ago

Staff Site Reliability Engineer, Security

Stord · Remote, United States, United States of America · Remote

1 month ago

Sr Staff Site Reliability Engineer (SRE)

Arrow · Ahmedabad, India +4

1 month ago

Staff Site Reliability Engineer

Zoox · Foster City, CA · Hybrid

2 months ago

Staff Site Reliability Engineer (Collaboration Engineering)

NBCUniversal · Orlando, FL, United States · Hybrid

2 months ago
Staff Site Reliability Engineer- Eng at UKG | Hiring.Camp