Hiring.Camp

Site Reliability Engineer IV

mtb

·

2 weeks ago

Salary
$140k – $233k
Location
Buffalo, NY, United States of America
Workplace
Hybrid
Type
Full-time
Department
Engineering
Source
Workday

Description

Overview

Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence, observability, testing, and automation across the Software Development Lifecycle. Leads complex initiatives, influences enterprise engineering standards, and partners with senior stakeholders to improve system stability, resiliency, performance, and operational maturity. Serves as a mentor and technical leader for less experienced engineers across Technology.

Primary Responsibilities

  • Accountable for defining and driving service reliability standards, including SLOs, SLAs, SLIs, and error budgets across platforms.
  • Design and implement highly available, fault-tolerant architectures aligned with enterprise scalability and resiliency requirements.
  • Lead initiatives to improve system reliability, availability, performance, and operational excellence through automation and engineering best practices.
  • Develop and promote observability strategies leveraging logging, monitoring, alerting, distributed tracing, OpenTelemetry (OTel), Dynatrace, dashboards, and telemetry analytics.
  • Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.
  • Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.
  • Lead incident management practices, including detection, response, escalation, recovery, and coordination of high-severity production events.
  • Drive problem management and Root Cause Analysis (RCA) activities to prevent systemic issues and ensure corrective actions are implemented.
  • Lead automation initiatives for self-healing systems, operational workflows, deployments, recovery procedures, and reliability controls.
  • Partner with development teams to build reliable, observable, and scalable services throughout the Software Development Lifecycle (SDLC).
  • Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.
  • Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.
  • Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for infrastructure provisioning, configuration management, and environment standardization.
  • Support and optimize cloud environments, including Microsoft Azure services, deployment automation, scaling strategies, and application lifecycle management.
  • Utilize cloud-native monitoring and operational tools to improve platform visibility, reliability, and performance.
  • Serve as a technical authority for performance engineering, resilience, capacity planning, and workload optimization.
  • Drive production readiness practices, including performance testing, resiliency testing, failover validation, disaster recovery preparedness, and operational readiness reviews.
  • Review architectural designs and technical roadmaps, providing recommendations to improve reliability, scalability, resiliency, and operational efficiency.
  • Lead cross-team reliability improvement initiatives and influence enterprise engineering standards.
  • Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.
  • Partner with development, infrastructure, cybersecurity, architecture, and support teams to identify risks, drive continuous improvement, and optimize platform performance.
  • Participate in and lead post-incident reviews, ensuring actionable outcomes and measurable improvements.
  • Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
  • Present reliability initiatives, operational metrics, and engineering recommendations in architecture reviews, technical forums, and leadership discussions.
  • Mentor engineers on reliability engineering, observability, cloud engineering, automation, and operational best practices.
  • Engage with stakeholders to identify risks, dependencies, and optimization opportunities.
  • Ensure adherence to risk and regulatory standards and escalate issues when needed.
  • Promote an environment that supports a culture of belonging and reflects the M&T Bank brand.
  • Maintain internal control standards, including timely implementation of audit findings, regulatory requirements, and compliance expectations.
  • Complete other related duties as assigned.

Scope of Responsibilities

Applies expert-level Site Reliability Engineering practices across multiple platforms. Drives enterprise-wide reliability improvements and influences technical direction without direct authority. Provides technical leadership, mentorship, and guidance to engineers and project teams.

Supervisory/Managerial Responsibilities

No supervisory responsibilities.

Education and Experience Required

Associate’s degree and a minimum of 9 years’ systems analysis and/or application development work experience or Bachelor’s degree and a minimum of 7 years’ systems analysis and/or application development work experience. In lieu of a degree, a combined minimum of 11 years’ education and/or relevant work experience, including a minimum of 7 years’ systems analysis and/or application development work experience.

Expert experience in system design, reliability engineering, and production operations.

Advanced proficiency in at least one programming or scripting language.

Education and Experience Preferred

  • Experience with observability and incident management tooling.
  • Experience with cloud platforms such as AWS or Azure.
  • Strong understanding of CI/CD, DevOps, and SDLC practices.
  • Experience defining and implementing SLO/SLI frameworks.
  • Experience in regulated environments such as financial services.
  • Strong communication and stakeholder management skills.
  • Experience with Infrastructure as Code (IaC), including Terraform.
  • Experience with monitoring and observability platforms such as Dynatrace, OpenTelemetry, Azure Monitor, Application Insights, or similar tools.
  • Experience with automated testing, deployment automation, and reliability engineering practices.
  • Knowledge of capacity planning, resiliency testing, disaster recovery, and high-availability architectures.
  • Experience working in Agile and DevOps operating models.
  • Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies.
  • Industry certifications related to Cloud Engineering, Azure, AWS, Terraform, or Site Reliability Engineering preferred.

#LI-JB3

M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

Location

Buffalo, New York, United States of America

Skills

PythonAWSAzureTerraformCI/CDCybersecurityAgileDevOpsCompliance

Similar Jobs

30

Site Reliability Engineer

Unum · Atlanta, Georgia, USA, United States of America

Yesterday

Site Reliability Engineer

Kiongroup · Kraków, Poland · Hybrid

Yesterday

Site Reliability Engineer

Thales · Noida Berger Tower, India · Hybrid

Yesterday

Site Reliability Engineer

LSEG · IND-BLR-Divyasree Technopolis, India

Yesterday

Site Reliability Engineer

Experian · Cyberjaya, Selangor, Malaysia · Hybrid

2 days ago

Site Reliability Engineer

Cisco · USA-SAN FRANCISCO, United States of America +11 · Remote

2 days ago

Site Reliability Engineer

LSEG · ROU-Bucharest-Iuliu Maniu Boulevard, Romania

2 days ago

Site Reliability Engineer

Rbs · Bengaluru, India

2 days ago

Site Reliability Engineer

Huntington · Easton Ops Cols C Oh, United States of America +10 · Onsite

5 days ago

Site Reliability Engineer

Westpac Group · Sydney, NSW, Australia

5 days ago

Site Reliability Engineer

Professional Kyndryl · PELML Lima (PELML) La Molina, Peru · Remote

6 days ago

Site Reliability Engineer

"Arena Intelligence, Inc." · Bay Area · Hybrid

6 days ago

Site Reliability Engineer

Damia Group · Lisbon, Coimbra, Braga · Remote, Hybrid, Onsite

6 days ago

Site Reliability Engineer

Accenture · Taguig, Uptown Bonifacio Tower 3, Philippines

6 days ago

Site Reliability Engineer

NielsenIQ · Mexico City, MEX, Mexico

6 days ago

Site Reliability Engineer

Tinybird · Spain · Remote

6 days ago

Site Reliability Engineer

Cisco · USA-RESEARCH TRIANGLE PARK, United States of America +4 · Hybrid

1 week ago

Site Reliability Engineer

Miqdigital · Bengaluru, India

1 week ago

Site Reliability Engineer

Forwardnetworks · Santa Clara, CA +1

1 week ago

Site Reliability Engineer

Bosonai · Toronto · Remote

1 week ago

Site Reliability Engineer

Experian · Cyberjaya, Selangor, Malaysia · Hybrid

1 week ago

Site Reliability Engineer

CRC Careers · Charlotte NC - 600 S Tryon St., United States of America

1 week ago

Site Reliability Engineer

Citi Bank · 5900 HURONTARIO STREET MISSISSAUGA, Canada · Hybrid

1 week ago

Site Reliability Engineer

HostPapa · Remote

1 week ago

Site Reliability Engineer

citibank · Mississauga, ON,CA, CA · Hybrid

1 week ago

Site Reliability Engineer

CRC Careers · CRC - Charlotte, NC 600 S. Tryon St., United States of America

1 week ago

Site Reliability Engineer

LSEG · Taipei - Nan Shan Plaza, Taiwan

1 week ago

Site Reliability Engineer

Careers Home · Zagreb (Croatia) +4 · Hybrid

1 week ago

Site Reliability Engineer

Braiins · Braiins · Hybrid

1 week ago

Site Reliability Engineer

Acronis · Serbia +2 · Remote

1 week ago
Site Reliability Engineer IV at mtb • $140k – $233k | Hiring.Camp