Hiring.Camp

Sr. Observability Engineer II

Nextgen

·

Today

Location
Bangalore, India
Workplace
Remote
Type
Full-time
Department
Engineering
Experience
7+ years
Source
Workday

Description

Job Description:

The Sr. Observability Engineer II is a specialized engineering role serving as the subject matter expert for the Dynatrace observability platform. The primary purpose of this role is to mature the firm’s proactive monitoring capabilities by focusing on signal quality engineering and operational noise reduction. This position functions as a key technical liaison between Hosting Operations, Site Reliability Engineering (SRE), and other engineering teams to ensure the stability and reliability of our client-facing healthcare SaaS environment.

  • Engineer and maintain the Dynatrace multi-tenant environment configuration, including management zones, alerting profiles, metric and custom events, tagging strategies, and access boundaries.
  • Own the end-to-end alert lifecycle, defining ownership, naming standards, severity mapping, alert disposition taxonomies, and production-readiness criteria.
  • Systematically eliminate alert fatigue across top problem patterns utilizing advanced tuning, thresholds, correlation, suppression, and alert retirement, while maintaining robust coverage for real, client-impacting issues.
  • Design and implement synthetic monitors and business-centric service-health models for critical client journeys and access paths, ensuring degradation is detected in service and business terms rather than raw infrastructure metrics.
  • Author and standardize Dynatrace Query Language (DQL) queries, notebooks, and dashboards for operations, leadership, and reliability reporting, establishing best-practice development standards for the team.
  • Collaborate with tools administrators to manage observability configurations using configuration-as-code practices, ensuring monitoring setups are versioned, peer-reviewed, and reproducible.
  • Partner with Tools, Automation and SRE teams to enhance alert-to-ticket enrichment (such as Salesforce ITSM integrations) with context and probable cause, minimizing triage times and enabling automated remediation.
  • Partner with SQL, platform, OS, and API specialists to resolve monitoring blind spots, ensure domain-specific alerts are meaningful, and verify that alerts are backed by robust runbooks.
  • Support service-level (P1–P4) measurements by establishing reporting baselines for corporate accountability, ensuring noise-reduction changes are paired with strict operational guardrails.
  • Maintain comprehensive reference materials, standard operating procedures, and observability standards documentation in Confluence to support continuous team enablement and operational onboarding.
  • Perform other duties that support the overall objective of the position.

Education Required:

  • Bachelor’s Degree in Computer Science, Information Technology, or a related technical field.
  • Or, any combination of education and experience which would provide the required qualifications for the position.

Experience Required:

  • 7+ years of experience in observability, application performance monitoring (APM), monitoring engineering, site reliability engineering (SRE), or a closely related technology infrastructure field.
  • Demonstrable hands-on expertise with the Dynatrace platform, including Davis AI, DQL, dashboards, notebooks, management zones, alerting profiles, metric/custom events, and synthetic monitors.
  • Proven track record of designing, configuring, and tuning alerting systems at enterprise scale to successfully reduce noise while maintaining robust detection capabilities.
  • Experience integrating monitoring systems with enterprise ITSM/ticketing systems (such as Salesforce) for automated ticket routing and enrichment.
  • Experience working within highly regulated hosting environments (e.g., healthcare, HIPAA/HITRUST, SOC 2, or ISO 27001).

License/Certification Required:

  • Dynatrace certification (or commitment to obtain certification within 1 year of hire).

Knowledge, Skills & Abilities:

  • Knowledge of: Strong working knowledge of AWS infrastructure, Windows and Linux operating systems, and SQL Server database concepts. Advanced understanding of cloud-native observability frameworks, APM tools (specifically Dynatrace), and alert lifecycle management. Strong command of cloud computing (AWS), systems integration, database operations, and scripting (e.g., Python or PowerShell) to support automation.
  • Skill in: Excellent problem-solving capabilities. Strong written and verbal communication skills. Highly organized, detail-oriented, and self-driven, with a strong ownership mindset toward maintaining signal quality, reducing operational toil, and defending client service-level agreements.
  • Ability to: Proven ability to translate complex technical infrastructure metrics into direct service and business impacts. Ability to collaborate effectively across operations command analysts, SRE partners, database specialists, and leadership.

The company has reviewed this job description to ensure that essential functions and basic duties have been included. It is intended to provide guidelines for job expectations and the employee's ability to perform the position described. It is not intended to be construed as an exhaustive list of all functions, responsibilities, skills and abilities. Additional functions and requirements may be assigned by supervisors as deemed appropriate. This document does not represent a contract of employment, and the company reserves the right to change this job description and/or assign tasks for the employee to perform, as the company may deem appropriate.

NextGen Healthcare is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Skills

PythonAWSLinuxSQLSQL ServerConfluenceSalesforceSOCSRESOC 2HIPAAISO 27001

Similar Jobs

30

Senior Observability Engineer

Ci · CDA ON Head Office - 15 York, Canada · Remote, Hybrid

1 week ago

Senior Observability Engineer

Saxobank · Headquarters, Denmark · Remote, Hybrid

2 weeks ago

Senior Observability Engineer

Rbc · 250 NICOLLET MALL:MINNEAPOLIS, United States of America

2 months ago

Senior Observability Engineer

modernatx · POL - Mazowieckie - Warsaw - MESH Rondo Ignacego Daszynskiego 1, Poland +1 · Remote, Hybrid

2 months ago

Senior Observability Engineer

Rent the Runway · Galway, Ireland +1

2 months ago

Senior Observability Engineer

M&G · Pune, India

2 months ago

Sr. Observability Engineer

Micron Technology · Remote

9 months ago

Sr. Observability Engineer

Micron · Hyderabad - Phoenix Aquila, India

9 months ago

Senior Site Reliability Engineer - Observability

Westpac Group · Sydney, NSW, Australia

Yesterday

Senior Backend Software Engineer, AI Observability & Evals Platform (LangSmith)

LangChain · Boston, MA +1 · Onsite

4 days ago

Senior Fullstack Engineer, AI Observability & Evals Platform

LangChain · Boston, MA +1 · Onsite

4 days ago

Senior Observability / DevOps Engineer

Miratech · Kyiv, Kyiv city, Ukraine · Remote

4 days ago

Senior AWS DevOps Engineer - AWS, Kubernetes, HCP, CI/CD, Observability with AI-focus (REMOTE)

Koniag Government Services · Remote

5 days ago

Senior Systems Software Engineer, Observability and Telemetry Platform

Nvidia · US, CA, Santa Clara, United States of America +1 · Remote

5 days ago

Senior Systems Software Engineer, Observability and Telemetry Platform

Nvidia · Santa Clara, CA,US, US +1 · Remote

5 days ago

Observability Engineer I, II, III, or Senior

Tri-State Generation and Transmission Association, Inc. · Frederick, CO, United States, US

5 days ago

Senior Software Engineer (Observability)

Canva · Brisbane, QLD, Australia · Hybrid

1 week ago

Senior Software Engineer (Observability)

Canva · Melbourne, VIC, Australia · Hybrid

1 week ago

Senior Infrastructure Engineer - AI, Automation, Observability and Monitoring

Nvidia · US, CA, Santa Clara, United States of America

1 week ago

Senior Infrastructure Engineer - AI, Automation, Observability and Monitoring

Nvidia · Santa Clara, CA,US, US

1 week ago

Senior Site Reliability Engineer - Linux Systems & Application Observability

tastylive · Chicago, Illinois +1 · Remote, Hybrid, Onsite

1 week ago

Senior Site Reliability Engineer - Linux Systems & Application Observability

tastytrade · Chicago, Illinois +1 · Remote, Hybrid, Onsite

1 week ago

Senior Software Engineer, Agentic AI and Observability

Nvidia · US, CA, Santa Clara, United States of America

1 week ago

Senior Software Engineer, Agentic AI and Observability

Nvidia · Santa Clara, CA,US, US

1 week ago

Senior Engineer – Observability Engineering

Levi Strauss & Co. Careers · Mexico, D.F., Mexico +5 · Remote

1 week ago

Senior Observability & Monitoring Engineer

Fiserv is the global leader · Sao Paulo - Paulista, Brazil · Onsite

1 week ago

Senior AI Infrastructure Engineer, Observability

Firmus Technologies · Singapore

1 week ago

Senior Infrastructure Engineer - Network Observability

Truist · Atlanta GA - 303 Peachtree Center Avenue - Garden Offices, United States of America +2

1 week ago

Senior Platform Engineer - Observability

Capgroup · Charlotte, United States of America

1 week ago

Senior Enterprise Observability Operations Engineer

Careers Home · ABC Manila Office, Philippines · Onsite

1 week ago