Hiring.Camp

Staff Site Reliability Engineer

Servicetitan

·

5 days ago

Location
India Bengaluru, Karnataka
Type
Full-time
Department
Engineering
Seniority
Senior
Source
Workday

Description

Ready to be a Titan?

The Staff Site Reliability Engineer will be a key player in managing, optimizing, and ensuring the reliability and scalability of our SQL Server and PostgreSQL databases both in the cloud and on-premises. The ideal candidate will have extensive experience with Azure and AWS platforms, with a strong preference for Azure expertise. You will work closely with our development and operations teams to drive improvements in database performance, automate processes, and implement robust backup and recovery procedures.
 

What you'll do:

  • Own the architecture, design, deployment, and lifecycle management of SQL Server and PostgreSQL databases across Azure, AWS, and on-prem environments.
  • Lead database design reviews, schema governance, indexing strategy, and query optimization to ensure performance and scalability.
  • Manage database security, including access controls, encryption (at rest and in transit), auditing, and compliance best practices.
  • Design and maintain high availability (HA) and disaster recovery (DR) architectures (Always On, replication, failover clusters, logical/physical replication, etc.), defining and enforcing RTO/RPO objectives.
  • Implement and manage backup strategies, validation testing, and recovery procedures.
  • Perform proactive database performance tuning, capacity planning, and workload optimization for mission-critical systems.
  • Own patching, upgrades, migrations (cloud/on-prem), and version lifecycle management.
  • Develop automation for DBA operations (provisioning, patching, backups, health checks, failover testing) using scripting (PowerShell, Bash, Python) and Infrastructure as Code.
  • Collaborate with engineering teams to optimize data models, troubleshoot production incidents, and resolve performance bottlenecks.
  • Establish database monitoring standards and implement observability solutions using tools such as Datadog, Grafana, ELK, and Prometheus.
  • Participate in incident response, root cause analysis (RCA), and postmortem improvements related to database systems.
  • Contribute to CI/CD processes for database deployments, schema changes, and migration pipelines.
  • Define and document database operational standards, runbooks, and best practices.

What you'll bring:
• 8+ years of database engineering/administration experience, including demonstrated ownership of platform-level database strategy and reliability across multiple teams or business-critical systems.
• Deep expertise in: Performance tuning (query plans, indexing strategies, locking/blocking analysis, deadlock resolution): High availability and disaster recovery configurations; Backup/restore strategies and validation; Database security and hardening
• Strong experience operating databases in Azure and AWS, including managed services (Azure SQL, RDS, etc.) and self-managed deployments.
• Proven experience with database migrations (on-prem to cloud, version upgrades, cross-platform migrations).
• Strong scripting skills (PowerShell, Bash, Python) to automate DBA workflows.
• Experience implementing monitoring, alerting, and observability for database systems.
• Solid understanding of reliability engineering principles (SLIs/SLOs) as they apply to database systems.
• Experience with Infrastructure as Code and containerized database environments (Kubernetes, Docker) is a plus.
• Familiarity with CI/CD pipelines for database schema deployments (GitHub Actions, Azure DevOps, TeamCity, etc.).
• Strong troubleshooting skills, attention to detail, and ability to manage high-impact production systems.

Be Human With Us: 

Being human isn’t about checking every box on a list. It’s about the experiences we have, people we meet, and the perspectives we share. So, if you have the skills but are hesitant to apply because of your background, apply anyway. We need amazing people like you to help us challenge the conventional and think differently about the problems that we’re solving. We’re in this together. Come be human, with us. 

Use of AI Technology:

We use technology, including automated and AI-assisted tools, to support certain aspects of our recruitment process. These tools are designed to improve efficiency and enhance the candidate experience. AI tools are not used to make hiring decisions; all hiring decisions are made by our hiring teams.

At ServiceTitan, we celebrate individuality and uniqueness. We believe that the convergence of fresh perspectives and experiences from all walks of life is what makes our product and culture so great. We do not discriminate against employees based on race, color, religion, sex, national origin, gender identity or expression, age, disability, sexual orientation, or any other characteristic protected by applicable laws. 

Skills

PythonAWSAzureDockerKubernetesCI/CDSQLPostgreSQLSQL ServerGitHubDevOpsCompliance

Similar Jobs

30

Staff Site Reliability Engineer

ServiceNow · Dublin, Ireland · Hybrid

3 days ago

Staff Site Reliability Engineer

Ionq · Santa Clara, California, United States

3 days ago

Senior/Staff Site Reliability Engineer

Pinecone · Tel Aviv

4 days ago

Sr/Staff Site Reliability Engineer, Consumer Apps

Attain · Chicago, IL +1 · Remote, Hybrid, Onsite

4 days ago

Staff Site Reliability Engineer

SimSpace Corporation · Remote - U.S. · Remote

5 days ago

Senior Staff Site Reliability Engineer

Pingidentity · US - Remote +1 · Remote

6 days ago

Staff Site Reliability Engineer

Levi Strauss & Co. Careers · Bengaluru, India

1 week ago

Staff Site Reliability Engineer

Tenex · Remote, USA · Remote

1 week ago

Sr. Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)

Zscaler · Hyderabad, IND

1 week ago

Staff Site Reliability Operations Engineer

Calix · Bangalore, India · Remote, Onsite

1 week ago

Staff Site Reliability Engineer

Stryker is one of the · Haryana, Gurugram International Techpark, Block I Phase 1 Floors G, 3, 4, 5, India · Hybrid

2 weeks ago

Staff Site Reliability Engineer

Chamberlain · Oak Brook, United States of America

2 weeks ago

Staff Site Reliability Engineer

Stackblitz · Remote · Remote

2 weeks ago

Staff Platform Engineer / Staff Site Reliability Engineer

Andurilindustries · Sydney, New South Wales, Australia

2 weeks ago

Staff Site Reliability Engineer - Paze

earlywarningservices · San Francisco, United States of America +2 · Hybrid

2 weeks ago

Staff Site Reliability Operations Engineer

Calix · Bangalore, India · Remote

3 weeks ago

Staff Site Reliability Engineer

The Onset Jobs Marketplace · AU

3 weeks ago

Staff Site Reliability Engineer- Eng

UKG · Lowell, MA,US, US

3 weeks ago

Staff Site Reliability Engineer - Volcano

Kong · United States · Remote

4 weeks ago

Staff Site Reliability Operations Engineer

Calix will come from · Remote - USA, United States of America · Remote

1 month ago

Staff Site Reliability Engineer, Security

Stord · Remote, United States, United States of America · Remote

1 month ago

Sr Staff Site Reliability Engineer (SRE)

Arrow · Ahmedabad, India +4

1 month ago

Staff Site Reliability Engineer

Zoox · Foster City, CA · Hybrid

1 month ago

Staff Site Reliability Engineer

Snyk · Portugal - Lisbon Office · Hybrid

1 month ago

Staff Site Reliability Engineer

Pingidentity · USA - Remote +1 · Remote

1 month ago

Staff Site Reliability Engineer — Project Volcano

Kong · United States · Remote

1 month ago

Staff Site Reliability Engineer (Collaboration Engineering)

NBCUniversal · Orlando, FL, United States · Hybrid

1 month ago

Senior Staff Site Reliability Engineer

Ironcladhq · San Francisco +2 · Hybrid

1 month ago

Sr Staff Site Reliability Engineer

Archer56 · San Jose, California, United States

1 month ago

Senior Staff Site Reliability Engineer

Hivewatch · El Segundo, CA +1

1 month ago
Staff Site Reliability Engineer at Servicetitan | Hiring.Camp