Hiring.Camp

Software Engineer, Reliability (SRE)

Veeamsoftware

·

Today

Location
Bangalore, India
Workplace
Hybrid
Department
VDC - Engineering 1009422
Experience
2+ years
Source
Greenhouse

Description

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role :

We are looking for a Senior Software Engineer, Reliability, you will serve as a hands-on technical leader within the SRE team, guiding senior engineers, influencing product development teams, and ensuring the systems we operate are built to be reliable, scalable, and observable from the ground up.

You will drive strategic initiatives, mentor others in the practice of SRE, and help define architectural best practices across our platform. This role is pivotal in aligning teams, enforcing high standards, and scaling SRE principles globally within Veeam.

What You'll Do :

Reliability Engineering & Resilience :

- Design and evolve infrastructure to be highly available, fault tolerant, and scalable across public clouds (initially Azure, with future expansion plans to other providers).

- Establish and maintain SLIs, SLOs, and error budgets that define and enforce reliability objectives.

- Lead incident response, analysis, blameless postmortems, and sharing sessions in order to maximize learning across our entire engineering team and driving changes to the entire socio-technical engineering system.

Observability & Operational Excellence :

- Drive adoption of deep observability practices, ensuring telemetry, logs, metrics, and tracing are comprehensive and actionable.

- Develop automation and self-healing tools to reduce toil and support Veeam's fleet management strategy.

- Participate in on-call rotations and lead operational excellence across the stack.

Engineering at Scale :

- Contribute to infrastructure as code (IaC), CI/CD systems, deployment automation, and scalable config management.

- Integrate and extend monitoring and chaos engineering tools to validate reliability assumptions under load and failure conditions.

- Implement testing strategies, canary deployments, and release validation pipelines to protect production environments and allow teams to safely deliver new features as quickly as possible.

Collaboration & Culture :

- Embed within product and platform teams to champion reliability from design through delivery.

- Contribute to a learning culture focused on continuous improvement and proactive risk management.

- Mentor engineers and advocate for DevOps/SRE best practices across global teams.

What You'll Bring :

- 5+ years of hands-on experience in a Software Engineering role with at least 2 years in Site Reliability, Platform Engineering, or similar.

- Deep experience building systems on public cloud providers (Azure preferred)

- Strong programming skills in JS, Node, Typescript, Go, Java, C#, or similar.

- Proven track record in delivering monitoring, alerting, and observability tooling (e.g., Prometheus, Grafana, OpenTelemetry).

- Experience with IaC tools like Terraform/Pulumi, and container orchestration (e.g., Kubernetes).

- Solid understanding of distributed systems, cloud networking, and cloud-native system design.

- Excellent communication and collaboration skills across geographies and disciplines.

Bonus Skills :

- Experience working on large-scale B2B SaaS platforms.

- Background in chaos engineering, resilience testing, performance testing, load testing, or incident learning programs.

- Familiarity with compliance frameworks (e.g., ISO, SOC 2, GDPR, FEDRAMP/CMMC).

What You'll Get :

- 18 paid vacation days, plus 4 extra global VeeaMe Days for self-care and 24 paid volunteer hours annually through Veeam Cares

Private medical coverage for you and up to four dependents

- Life, accident, and disability insurance with enhanced coverage

- Annual flexible wellbeing allowance for physical and mental wellness

- Free confidential counselling and coaching via Employee Assistance Program (EAP), including legal and financial advice

- Meal, fuel, and transportation benefits based on work arrangement

- Daycare reimbursement and safe cab facility for eligible employees

- Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O'Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning

Please note : If the applicant is permanently located outside India, Veeam reserves the right to decline the application.

#LI-SK2

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential. Personal data collected during the recruitment process will be processed in accordance with ourRecruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

 

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

Skills

TypeScriptJavaAzureKubernetesTerraformCI/CDSOCDevOpsSRERisk ManagementComplianceSOC 2GDPRGo

Similar Jobs

30

Software Reliability Engineer

Tmobile · GA-Atlanta Ravinia Office, United States of America

1 month ago

Software Engineer, Reliability

Openai · San Francisco

9 months ago

Open to Jobshare - Sr Lead Software Engineer - Site Reliability Engineer, Python & Infrastructure management - Part time/Jobshare

JPMorgan Chase · LONDON, LONDON, United Kingdom, GB

Today

Open to Jobshare - Sr Lead Software Engineer - Site Reliability Engineer, Python & Infrastructure management - Part time/Jobshare

JP Morgan Chase · LONDON, LONDON, United Kingdom, GB

Today

Software Development Engineer, Region Reliability

Amazon

Today

Software Engineer, Reliability (SRE)

Veeamsoftware · Pune, India · Hybrid

Today

Software Engineer III - AI/ML Platform Reliability

JPMorgan Chase · GLASGOW, LANARKSHIRE, United Kingdom, GB

Yesterday

Software Engineer III - AI/ML Platform Reliability

JP Morgan Chase · GLASGOW, LANARKSHIRE, United Kingdom, GB

Yesterday

Software Engineer III, Site Reliability Engineering

JPMorgan Chase · Tokyo-To, Japan, JP

2 days ago

Software Engineer III, Site Reliability Engineering

JP Morgan Chase · Tokyo-To, Japan, JP

2 days ago

Software Engineer, Site Reliability Engineering

bet365 · Manchester, England, United Kingdom · Hybrid

5 days ago

Software Engineer, Site Reliability Engineering

bet365 · Stoke-on-Trent, England, United Kingdom · Hybrid

5 days ago

Lead Software Engineer - AI Platform Reliability

JPMorgan Chase · Seattle, WA, United States, US

1 week ago

Lead Software Engineer - AI Platform Reliability

JP Morgan Chase · Seattle, WA, United States, US

1 week ago

Senior Software Engineer, Backend (Reliability Platform)

Affirm · Remote Canada · Remote

1 week ago

Senior Software Engineer, Backend (Reliability Platform)

Affirm · Remote US · Remote

1 week ago

Associate, Software Production Management & Reliability Engineer

Ms · Otemachi Financial City, Japan

1 week ago

Associate, Software Production Management & Reliability Engineer

Ms · Otemachi Financial City, Japan

1 week ago

Associate, Software Production Management & Reliability Engineer

Morgan Stanley · Tokyo,JP, JP

1 week ago

Staff Software Engineer, Reliability

Metropolis · Bengaluru, Karnataka, India

1 week ago

Software/Site Reliability Engineer - FedRAMP

"Tenable, Inc." · US - Remote - California - Bay Area · Remote

1 week ago

Sr Software Engineer - Reliability Engineering

Cox brings together the world · North Hills, NY - 3400 New Hyde Park Rd, United States of America · Remote, Hybrid

1 week ago

Site Reliability Engineer - Software Ops and Scaling , One Material Handling System - Software, Controls and Science

Amazon

2 weeks ago

Java AI Lead Software Engineer, Automation and Reliability

JPMorgan Chase · OH, United States, US

2 weeks ago

Java AI Lead Software Engineer, Automation and Reliability

JP Morgan Chase · OH, United States, US

2 weeks ago

Senior Lead Software Engineer - LLM Ops Platform Reliability

JPMorgan Chase · GLASGOW, LANARKSHIRE, United Kingdom, GB

2 weeks ago

Senior Lead Software Engineer - LLM Ops Platform Reliability

JP Morgan Chase · GLASGOW, LANARKSHIRE, United Kingdom, GB

2 weeks ago

Production Engineer, Site Reliability (Application Software)

Spacex · Hawthorne, CA +1

2 weeks ago

Software Engineer Lead - Site Reliability Engineering Center

PNC Bank · The Tower at PNC Plaza (PAA86), United States of America +5 · Onsite

2 weeks ago

Software Engineer, Site Reliability Engineering (Application Software)

Spacex · Hawthorne, CA +1

2 weeks ago