Hiring.Camp

Manager, Solution Engineering(Site Reliability Engineering)

Westernunion

·

Today

Location
Pune - Business Bay, India
Workplace
Hybrid
Type
Full-time
Department
Engineering
Seniority
Manager
Closing date
Today
Source
Workday

Description

Role Responsibilities


At Western Union, reliability is a product feature. Our customers depend on our platforms to move money securely and seamlessly across the globe, and every second of availability matters.

We're looking for an Engineering Manager to lead our Site Reliability Engineering organization and help shape the future of reliability, resilience, and operational excellence across our most critical platforms that support millions of transactions every day. This is not a traditional production support role. You will lead teams that apply software engineering principles to operations, building automated, self-healing, highly observable systems feature that scale globally and operate with exceptional reliability and accelerate innovation through automation and AI-driven operational practices.

This role offers a unique opportunity to influence enterprise-wide reliability strategy, modernize operational capabilities, and build a high-performing engineering culture across globally distributed teams supporting highly regulated financial services and compliance platforms.

You will also –

• Lead and mentor Site Reliability Engineering, Platform Engineering, and Production Support teams.

• Establish reliability, scalability, and operational excellence standards across critical platforms.

• Foster a culture of accountability, continuous improvement, innovation, and engineering excellence.

• Drive adoption of modern SRE practices, automation frameworks, and AI-assisted operations.

Reliability & Platform Engineering

• Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.

• Improve system reliability, availability, resilience, scalability, and performance.

• Lead capacity planning, disaster recovery, and business continuity initiatives.

• Design and implement proactive monitoring, observability, and anomaly detection capabilities.

Production Operations & Service Management

• Own and evolve the L2/L3 production support operating model.

• Lead and Support Incident, Problem, Change, and Release Management processes.

• Drive rapid incident response, service restoration, root cause analysis, and corrective actions.

• Ensure compliance with established SLAs, SLOs, operational controls, and audit requirements.

Automation & Operational Excellence

• Reduce operational toil through automation and self-healing capabilities.

• Identify and eliminate manual processes through engineering improvements and workflow optimization.

• Implement AI-powered operational insights and incident investigation capabilities.

• Drive continuous improvement initiatives focused on stability, efficiency, and cost optimization.

Stakeholder & Business Partnership

• Collaborate with Product, Engineering, Infrastructure, Security, Compliance, and Business Operations teams.

• Communicate service health, operational risks, and reliability metrics to senior leadership.

• Manage relationships with managed service providers and strategic technology partners.

• Support highly regulated financial services, payments, risk, fraud, and compliance platforms.

Role Requirements

Leadership & Experience

  • Bachelor’s degree in computer science, Information Systems, IT, or similar preferred
  • 12+ years in Software Engineering, Platform Engineering, DevOps, or SRE.
  • 5+ years managing SRE, Platform, or Production Support teams.
  • Experience leading 8-15+ engineers across multiple workstreams.
  • Proven experience supporting enterprise applications in Agile environments.
  • Experience leading global support teams supporting business-critical applications.
  • Strong stakeholder management and executive communication skills.

Site Reliability Engineering

  • Deep expertise in modern SRE practices, including SLOs, SLIs, Error Budgets, capacity planning, resiliency, and reliability engineering.
  • Demonstrated success by reducing operational toil through automation and building systems that are reliable by design.
  • Experience leading incident management, blameless postmortems, and continuous reliability improvement programs.

AWS Cloud , Platform & DevOps

  • Strong hands-on experience designing and operating cloud-native platforms on AWS, including EKS/ECS, EC2, IAM, API Gateway, S3, and related services.
  • Expertise in Kubernetes, Infrastructure as Code (Terraform), CI/CD, DevSecOps, and platform engineering.
  •  Passion for building self-service platforms and enabling teams to move faster with confidence.
  • Experience leveraging AI-assisted engineering and operations tools such as GitHub Copilot, Amazon Q, or equivalent AI tools to improve developer productivity and service reliability.
  • Proficiency in scripting languages (Python, Shell) and familiarity with automation tools (such as Ansible, Jenkins). 

Observability & Operational Excellence

  • Experience implementing observability strategies using tools such as  Dynatrace, Prometheus, Grafana, and OpenSearch.
  •  Strong understanding of incident response, disaster recovery, production readiness, performance engineering, and operational risk management.
  • Ability to establish meaningful reliability metrics and drive data-informed operational decisions.

Technical Foundation

  • Linux Administration, API troubleshooting , Distributed Systems , Networking fundamentals , Performance Engineering

Domain Expectations (Mandatory)

Preferred Industry Background

•             Experience operating mission-critical platforms in Financial Services, Payments,      Banking, FinTech, Fraud, Risk, or AML environments.

  • Familiarity with security, compliance, audit, and regulatory requirements associated with highly regulated, high-availability systems.

Nice-to-Have Skills

Advanced Reliability Engineering

  • Experience with Chaos Engineering, fault-injection testing, resilience validation, and large-scale performance benchmarking.
  • Experience driving reliability transformation or SRE maturity initiatives across enterprise organizations.

Work Shift

Western Union values in-person collaboration, problem solving, and ideation whenever possible. We believe this fosters common ways of working and supports how we execute initiatives for our customers. The expectation is to work from the office a minimum of three days a week.

BENEFITS AND OTHER DETAILS

Benefits

You will also have access to short-term incentives, multiple health insurance options, accident and life
insurance, and access to best-in-class development platforms, to name a few. Please see the benefits below specific to your country. If applicable, additional role-specific benefits will be mentioned during your interview process or in an offer of employment.

Your India specific benefits include:

  • Employees Provident Fund [EPF]
  • Gratuity Payment
  • Public holidays
  • Annual Leave, Sick leave, Compensatory leave, and Maternity / Paternity leave
  • Annual Health Checkup
  • Hospitalization Insurance Coverage (Mediclaim)
  • Group Life Insurance, Group Personal Accident Insurance Coverage, Business Travel Insurance
  • Cab Facility
  • Relocation Benefit

Other Details

As part of the application process, all applicants are required to take assessments. Western Union has partnered with a 3rd party provider to administer these tests. Applicants will need to provide their name and email address in order to process the assessments. If you have any questions, you may reach out to [email protected].

We are passionate about honouring our employee's identity and fostering a feeling of belonging. Our commitment is to provide an inclusive culture that celebrates the unique backgrounds and perspectives of our global teams while reflecting the communities we serve. We do not discriminate based on race, color, national origin, religion, political affiliation, sex (including pregnancy), sexual orientation, gender identity, age, disability, marital status, or veteran status. The company will provide accommodation to applicants, including those with disabilities, during the recruitment process, following applicable laws

#LI-MT #LI-Hybrid

Estimated Job Posting End Date:

08-21-2026

This application window is a good-faith estimate of the time that this posting will remain open. This posting will be promptly updated if the deadline is extended or the role is filled.

Skills

PythonAWSKubernetesTerraformAnsibleJenkinsCI/CDLinuxGitHubAgileDevOpsSREAMLRisk ManagementCompliance