- Location
- Bangalore, India
- Workplace
- Remote
- Type
- Full-time
- Department
- Engineering
- Seniority
- Lead
- Source
- Workday
Description
Job Description:
The Principal Database Engineer (DBE) – Cloud Platform Support – Global is a premier technical authority and operational anchor within the Global Cloud Platform Support organization. Operating primarily during the IST daytime shift (providing operational leadership and coverage for US overnight operations, with designated shift handover overlap), this role ensures 24/7 database stability, high availability, performance, and rapid incident resolution for NextGen’s cloud healthcare enterprise platforms.This position bridges deep database engineering with Site Reliability Engineering (SRE), DevOps, and 24/7 Support workflows. As a Principal DBE, you will drive hands-on actioning of complex SQL-based alerts, execute critical database performance tuning and maintenance requests, and serve as the highest-tier Subject Matter Expert (SME) and escalation authority for database functionality, query efficiency, and incident troubleshooting across global support operations.
Work Schedule & Global Operational Model:
- Shift Structure: IST Daytime hours providing senior technical leadership and coverage for US Overnight operations.
- Handover & Global Alignment: Structured daily overlap windows with US-based Cloud Support and Database Engineering leadership to ensure seamless operational continuity, bi-directional shift handoffs, and synchronized escalation governance.
Key Responsibilities:
Tier-3 Escalation SME & Advanced Alert Response:
- Serve as the global escalation authority and Subject Matter Expert (SME) for Support team members regarding database functionality, concurrency bottlenecks, and complex troubleshooting.
- Proactively monitor, triage, and action advanced SQL- and database-level alerts across enterprise production environments in AWS and GCP (leveraging Dynatrace, CloudWatch, GCP Operations, and native SQL diagnostic telemetry).
- Take command of high-severity production database incidents (Sev-1/Sev-2) during shift hours, driving rapid incident mitigation, deep diagnostic artifact capture, and post-incident stabilization.
- Analyze alert patterns, identify recurring systemic anomalies, and partner with SRE and engineering teams to eliminate alert fatigue and implement automated self-healing responses.
Database Performance Tuning & Workload Optimization:
- Act as the lead technical authority for executing, analyzing, and validating database performance tuning requests submitted by Support, Implementation, and Product Engineering teams.
- Conduct advanced execution plan analysis, dynamic management view/function (DMV/DMF) audits, index defragmentation/restructuring, and query recompilation strategies for mission-critical workloads.
- Diagnose live production concurrency challenges—including deep transaction blocking chains, deadlocks, lock escalation, latch contention, and tempdb bottlenecks—applying safe, decisive interventions.
- Provide consultative optimization guidance and query refactoring standards to application support engineers and developers to prevent persistent performance degradations.
Database Maintenance & Operational Resilience:
- Oversee and execute scheduled and complex database maintenance procedures during low-impact shift windows (including index maintenance, statistics updates, partition management, and DBCC CHECKDB integrity checks).
- Continuously validate nightly database backup integrity, snapshot consistency, and cross-region replication health to enforce stringent Recovery Point Objectives (RPO) and Recovery Time Objectives (RTO).
- Govern and perform high-impact operational database tasks, including point-in-time restores (PITR), schema object deployments, storage/volume expansions, failover drills, and configuration changes.
- Coordinate and support minor engine version upgrades, security patch cycles, and parameter group modifications across managed (RDS/Aurora, Cloud SQL) and self-managed cloud instances.
Engineering Bridge, Automation & Toil Reduction:
- Bridge deep database engineering principles with daily Support and SRE practices, creating reusable automation scripts (PowerShell, Python, Bash, T-SQL) to eliminate repetitive operational toil.
- Author enterprise-grade diagnostic runbooks, standard operating procedures (SOPs), and self-service troubleshooting tooling to elevate Tier-1 and Tier-2 support capabilities.
- Lead comprehensive post-incident reviews (PIR) and author detailed Root Cause Analysis (RCA) documentation following critical incidents, driving permanent preventative remediation.
- Partner with Platform Engineering to validate and test database release pipelines, schema migrations, and CI/CD database tooling.
Education Required:
- Bachelor’s Degree in Computer Science, Information Technology, Computer Applications (MCA), or a related technical discipline; or an equivalent combination of education, relevant certifications, and extensive professional experience.
- Or, any combination of education and experience which would provide the required qualifications for the position.
Experience Required:
- 10+ years of progressive experience in enterprise relational database administration and engineering, supporting mission-critical, high-availability 24/7 cloud production systems.
- Proven track record operating as a Principal or Lead Database Engineer/SME, comfortable operating autonomously during off-hours/overnight shifts with decisive incident response capabilities.
- Extensive hands-on experience designing, maintaining, and troubleshooting enterprise databases on AWS or GCP (RDS, Aurora, Cloud SQL, or EC2/Compute Engine IaaS deployments).
- Deep mastery of Microsoft SQL Server internals (SQL Server 2016 through 2022), including AlwaysOn Availability Groups, query optimizer behaviors, locking/latching mechanics, and storage engine architecture.
- Strong operational proficiency in PostgreSQL or equivalent open-source relational database engines in production environments.
- Proven expertise in query performance tuning, index optimization, wait-event analysis, and dynamic tracing in high-throughput enterprise SaaS applications.
- Solid experience utilizing APM and observability platforms (Dynatrace, Datadog, Prometheus/Grafana) for database metrics correlation and proactive alerting.
License/Certification Required:
- Microsoft Certified: Azure Database Administrator Associate or legacy MCSE: Data Management and Analytics.
- AWS Certified Database – Specialty or AWS Certified Solutions Architect – Professional.
- Google Cloud Professional Cloud Database Engineer.
- Practical familiarity with Infrastructure as Code (Terraform) and configuration automation (Ansible).
- Prior experience in Healthcare IT, HIPAA/HITRUST regulatory compliance, or EHR/PM clinical data systems.
- ITIL Foundation or SRE/DevOps certifications.
Knowledge, Skills & Abilities:
- Technical Knowledge
- Database Architecture & Internals: Expert knowledge of the SQL Server and PostgreSQL storage engines, transaction logging architecture, buffer pool management, ACID properties, isolation levels, and execution plan operators.
- High Availability & Business Continuity: Deep understanding of AlwaysOn AG multi-subnet failover, asynchronous/synchronous replica management, DR replication topologies, and cloud automated failover mechanisms.
- Cloud Infrastructure & Networking: Strong comprehension of cloud compute, provisioned IOPS storage types, VPC networking, security groups, and IAM least-privilege security controls.
- Core Skills
- Diagnostic & Tuning Mastery: Advanced capability to rapidly triage critical incidents, isolate root-cause wait stats, remediate parameter sniffing, and resolve severe blocking in real time.
- Scripting & Automation: High proficiency in PowerShell, Python, T-SQL, or Bash to script automated health checks, routine maintenance, and diagnostic log collection.
- Technical Mentorship & Communication: Exceptional written and verbal communication skills to conduct cross-continental handovers, author executive incident summaries, and mentor global support teams.
- Operational Competencies
- Autonomous Leadership: Exceptional ability to exercise sound architectural and operational judgment independently during off-hours shifts.
- Proactive Continuous Improvement: Continuous commitment to converting recurring operational issues into documented bug fixes, product enhancements, or automated remediations.
The company has reviewed this job description to ensure that essential functions and basic duties have been included. It is intended to provide guidelines for job expectations and the employee's ability to perform the position described. It is not intended to be construed as an exhaustive list of all functions, responsibilities, skills and abilities. Additional functions and requirements may be assigned by supervisors as deemed appropriate. This document does not represent a contract of employment, and the company reserves the right to change this job description and/or assign tasks for the employee to perform, as the company may deem appropriate.
NextGen Healthcare is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Skills
PythonAWSAzureGCPTerraformAnsibleCI/CDSQLPostgreSQLSQL ServerDevOpsSREEHRComplianceITILHIPAAAWS Certified