Hiring.Camp

Staff Site Reliability Engineer- Eng

UKG

·

Today

Location
Noida, UP,IN, IN
Type
Full-time
Department
Engineering
Seniority
Senior
Experience
10+ years
Source
Eightfold

Description

Why UKG:

At UKG, the work you do matters. The code you ship, the decisions you make, and the care you show a customer all add up to real impact. Today, tens of millions of workers start and end their days with our workforce operating platform. Helping people get paid, grow in their careers, and shape the future of their industries. That’s what we do.

We never stop learning. We never stop challenging the norm. We push for better, and we celebrate the wins along the way. Here, you’ll get flexibility that’s real, benefits you can count on, and a team that succeeds together. Because at UKG, your work matters—and so do you.

About the role

Site Reliability Engineers at UKG are team members that have a breadth of knowledge encompassing all aspects of service delivery. They develop software solutions to enhance, harden and support our service delivery processes. This can include building and managing CI/CD deployment pipelines, automated testing, capacity planning, performance analysis, monitoring, alerting, chaos engineering and auto remediation.

Site Reliability Engineers must have a passion for learning and evolving with current technology trends. They strive to innovate and are relentless in their pursuit of a flawless customer experience. They have an “automate everything” mindset, helping us bring value to our customers by deploying services with incredible speed, consistency and availability.

Primary/Essential Duties and Key Responsibilities:

  • Proficient in Splunk/ELK, and Datadog.
  • Experience with observability tools such as Prometheus/InfluxDB, and Grafana.
  • Possesses strong knowledge of at least one scripting language such as Python, Bash, Powershell or any other relevant languages.
  • Design, develop, and maintain observability tools and infrastructure.
  • Collaborate with other teams to ensure observability best practices are followed.
  • Develop and maintain dashboards and alerts for monitoring system health.
  • Troubleshoot and resolve issues related to observability tools and infrastructure.
  • Engage in and improve the lifecycle of services from conception to EOL, including: system design consulting, and capacity planning
  • Define and implement standards and best practices related to: System Architecture, Service delivery, metrics and the automation of operational tasks
  • Support services, product & engineering teams by providing common tooling and frameworks to deliver increased availability and improved incident response.
  • Improve system performance, application delivery and efficiency through automation, process refinement, postmortem reviews, and in-depth configuration analysis
  • Collaborate closely with engineering professionals within the organization to deliver reliable services
  • Identify and eliminate operational toil by treating operational challenges as a software engineering problem
  • Actively participate in incident response, including on-call responsibilities
  • Partner with stakeholders to influence and help drive the best possible technical and business outcomes
  • Guide junior team members and serve as a champion for Site Reliability Engineering
  • Engineering degree, or a related technical discipline, and 10+years of experience in SRE.
  • Experience coding in higher-level languages (e.g., Python, Javascript, C++, or Java)
  • Knowledge of Cloud based applications & Containerization Technologies
  • Demonstrated understanding of best practices in metric generation and collection, log aggregation pipelines, time-series databases, and distributed tracing
  • Ability to analyze current technology utilized and engineering practices within the company and develop steps and processes to improve and expand upon them
  • Working experience with industry standards like Terraform, Ansible.
  • (Experience, Education, Certification, License and Training)
  • Must have hands-on experience working within Engineering or Cloud.
  • Experience with public cloud platforms (e.g. GCP, AWS, Azure)
  • Experience in configuration and maintenance of applications & systems
  • infrastructure. Experience with distributed system design and architecture
  • Experience building and managing CI/CD Pipelines

Company Overview:

UKG is the Workforce Operating Platform that puts workforce understanding to work. With the world's largest collection of workforce insights, and people-first AI, our ability to reveal unseen ways to build trust, amplify productivity, and empower talent, is unmatched. It's this expertise that equips our customers with the intelligence to solve any challenge in any industry — because great organizations know their workforce is their competitive edge. Learn more at ukg.com.

UKG is proud to be an equal opportunity employer and is committed to promoting diversity and inclusion in the workplace, including the recruitment process.

Disability Accommodation in the Application and Interview Process

For individuals with disabilities that need additional assistance at any point in the application and interview process, please email [email protected]

Skills

PythonJavaScriptJavaAWSAzureGCPTerraformAnsibleCI/CDSplunkSRE

Similar Jobs

30

Staff Site Reliability Engineer

Trimble · Chennai, TN,IN, IN

Today

Staff Site Reliability Engineer (AI Platform)

Manychat · Barcelona, Spain +1 · Hybrid

Yesterday

Staff Site Reliability Engineer (AI Platform)

Manychat · Amsterdam, Netherlands +1

Yesterday

Staff Site Reliability Engineer

Guidewire · Kuala Lumpur Office, Malaysia · Hybrid

1 week ago

Staff Site Reliability Engineer

Pingidentity · UK - Remote +1 · Remote

1 week ago

Staff Site Reliability Engineer

Ionq · Santa Clara, California, United States

1 week ago

Sr/Staff Site Reliability Engineer, Consumer Apps

Attain · Chicago, IL +1 · Remote, Hybrid, Onsite

2 weeks ago

Staff Site Reliability Engineer

Servicetitan · India Bengaluru, Karnataka

2 weeks ago

Staff Site Reliability Engineer

SimSpace Corporation · Remote - U.S. · Remote

2 weeks ago

Staff Site Reliability Engineer

Levi Strauss & Co. Careers · Bengaluru, India

2 weeks ago

Staff Site Reliability Engineer

Tenex · Remote, USA · Remote

2 weeks ago

Sr. Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)

Zscaler · Hyderabad, IND

3 weeks ago

Staff Site Reliability Engineer

Stryker is one of the · Haryana, Gurugram International Techpark, Block I Phase 1 Floors G, 3, 4, 5, India · Hybrid

3 weeks ago

Staff Site Reliability Engineer

Chamberlain · Oak Brook, United States of America

3 weeks ago

Staff Site Reliability Engineer

Stackblitz · Remote · Remote

4 weeks ago

Staff Platform Engineer / Staff Site Reliability Engineer

Andurilindustries · Sydney, New South Wales, Australia

1 month ago

Staff Site Reliability Engineer - Paze

earlywarningservices · San Francisco, United States of America +2 · Hybrid

1 month ago

Staff Site Reliability Engineer

The Onset Jobs Marketplace · AU

1 month ago

Staff Site Reliability Engineer- Eng

UKG · Lowell, MA,US, US

1 month ago

Staff Site Reliability Engineer - Volcano

Kong · United States · Remote

1 month ago

Staff Site Reliability Operations Engineer

Calix will come from · Remote - USA, United States of America · Remote

1 month ago

Staff Site Reliability Engineer, Security

Stord · Remote, United States, United States of America · Remote

1 month ago

Sr Staff Site Reliability Engineer (SRE)

Arrow · Ahmedabad, India +4

1 month ago

Staff Site Reliability Engineer

Zoox · Foster City, CA · Hybrid

1 month ago

Staff Site Reliability Engineer (Collaboration Engineering)

NBCUniversal · Orlando, FL, United States · Hybrid

2 months ago

Senior Staff Site Reliability Engineer

Ironcladhq · San Francisco +2 · Hybrid

2 months ago

Sr Staff Site Reliability Engineer

Archer56 · San Jose, California, United States

2 months ago

Senior Staff Site Reliability Engineer

Hivewatch · El Segundo, CA +1

2 months ago

Staff Site Reliability Engineer - Site Experience

reddit · Remote - United Kingdom · Remote

2 months ago

Staff Site Reliability Engineer

Arcadia · Chennai, Tamil Nadu, India

2 months ago