Hiring.Camp

Site-Reliability Engineer, Application Operations

Carislifesciences

·

Yesterday

Location
Irving - HQ, United States of America
Type
Full-time
Department
Engineering
Experience
5+ years
Education
Bachelor
Source
Workday

Description

At Caris, we understand that cancer is an ugly word—a word no one wants to hear, but one that connects us all. That’s why we’re not just transforming cancer care—we’re changing lives.

 

We introduced precision medicine to the world and built an industry around the idea that every patient deserves answers as unique as their DNA. Backed by cutting-edge molecular science and AI, we ask ourselves every day: “What would I do if this patient were my mom?” That question drives everything we do.

 

But our mission doesn’t stop with cancer. We're pushing the frontiers of medicine and leading a revolution in healthcare—driven by innovation, compassion, and purpose.

 

Join us in our mission to improve the human condition across multiple diseases. If you're passionate about meaningful work and want to be part of something bigger than yourself, Caris is where your impact begins.

Position Summary

Caris Life Sciences is one of the largest precision-oncology platforms in the world, serving hundreds of thousands of molecular cases a year and growing at double-digit rates. Behind every case is a matched molecular, imaging, and clinical-outcomes data estate few organizations anywhere can rival, and the clinical software that drives the lab's instruments and processes, captures results, and delivers each patient's report. When that software degrades, patient care waits; keeping it healthy is this role's mission.

The Site-Reliability Engineer, Application Operations, works on the dedicated Application Site-Reliability Engineering (App-SRE) team. The work is site-reliability engineering across the clinical application portfolio, with real production telemetry and a direct hand in shaping the run-operate discipline.

This is a production-operations craft role, distinct from application feature development, operating under governed privileged access and segregation-of-duties discipline. The engineer serves as a primary responder in the on-call rotation; executes runbooks and approved maintenance scripts with audit-grade discipline; implements application instrumentation, SLO monitors, and error-budget tracking; supports releases, deployment health verification, and rollback execution; and contributes to post-incident reviews. The team runs automation-first: recurring manual work is engineering backlog, and the engineer progressively automates away the toil they encounter rather than absorbing it.

Frontier AI coding assistants are standard-issue tooling, with agentic workflows spanning runbook authoring, alert-quality analysis, and operational automation. Characterization tests and golden-master replay validation serve as executable evidence for regulated change. Delivery runs on CI/CD with application-level observability, spanning application-performance monitoring and production telemetry, and risk-based release governance aligned with FDA Computer Software Assurance guidance. Modernization of established systems is active engineering work, not deferred maintenance.

Reporting to the Director, Application Site-Reliability Engineering, this practitioner-level individual contributor operates production clinical applications subject to SOX financial controls and FDA regulatory requirements, working safely and accurately under established, documented procedures.

Location: Irving, Texas (Dallas–Fort Worth), on-site/hybrid, minimum three days per week on campus; co-located with Caris's laboratory and clinical operations.

Job Responsibilities

  • Participate in a scheduled on-call rotation as a primary responder for production incidents across the SOX- and FDA-regulated clinical application portfolio; execute triage, escalation, and initial remediation steps in accordance with documented runbooks and incident-response playbooks.

  • Execute approved runbooks and maintenance scripts in the production environment; document all privileged actions in compliance with SOX ITGCs and FDA audit-trail requirements.

  • Implement and maintain application-layer instrumentation, including dashboards, alert thresholds, and SLO monitors, in coordination with the centrally operated observability platform.

  • Support production releases and deployments: coordinate deployment health verification, monitor application behavior after deployment, and execute rollback procedures when required.

  • Participate in post-incident reviews: contribute timeline reconstructions, identify contributing factors, and track remediation action items through to closure.

  • Maintain audit-ready production-access logs, privileged-action records, and role-change documentation to support SOX ITGC and FDA regulatory audits.

  • Automate recurring operational work: convert repeated manual interventions, diagnostics, and maintenance procedures into scripted, reviewed, pipeline-executed automation under the team's change controls.

  • Track the manual-intervention rate for assigned services and drive it down over time by retiring runbook steps into automation.

  • Author and update runbook entries and known-issue documentation as operational knowledge is gained; contribute to continuous improvement of the operate discipline.

  • Collaborate with application engineering teams to gather context during incidents and to validate fixes deployed to the production environment.

  • Monitor application SLO attainment and error-budget consumption; escalate proactively when budgets are at risk.

  • Contribute to the ongoing development of on-call tooling, alert quality, and operational dashboards within the application-layer observability framework.

  • Work AI-first: use AI coding assistants and agentic workflows as daily practice in monitoring, triage, runbook work, and automation scripting, with review as the quality gate.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Systems, Software Engineering, or a closely related technical field, or equivalent practical experience.

  • 5+ years of professional experience in SRE, production operations, DevOps, or platform engineering in a production-support capacity.

  • 2+ years of direct, hands-on experience participating in an on-call rotation as a primary production responder for Tier 1 or business-critical systems.

  • Experience executing production runbooks, maintenance scripts, or change procedures in a production environment with documented privileged-action controls.

  • Experience implementing or maintaining application monitoring, alerting thresholds, or dashboards in a production observability platform.

  • Experience participating in post-incident review or blameless retrospective processes, including timeline reconstruction and corrective-action tracking.

  • Hands-on use of AI coding assistants for automation, scripting, or operational tooling.

Preferred Qualifications

  • Familiarity with SLO frameworks, error budgets, and associated alerting design patterns.

  • Experience reducing operational toil through scripting or automation in an SRE or production-operations setting.

  • Experience working in a SOX-controlled IT environment or a CLIA/CAP-regulated laboratory setting, including change-control ticketing, access-review processes, or audit-evidence collection.

  • Working knowledge of modern cloud-native observability at the application-instrumentation layer, including open standards for telemetry and tracing and application-performance-monitoring platforms.

  • Domain experience in clinical diagnostics, laboratory information systems, or digital health software.

  • Familiarity with SAST/DAST tooling and secure CI/CD pipelines, including pipeline-embedded security scanning.

  • Familiarity with deployment pipelines, container orchestration, and release-automation tooling from an operate and release-support perspective.

  • Certifications in cloud platforms or information-security disciplines relevant to production operations.

Physical Demands

  • Ability to sit, stand, and work at a computer for extended periods.

Training

  • All job-specific, safety, and compliance training is assigned based on the job functions associated with this employee.

Other

  • This role includes participation in a scheduled on-call rotation with required after-hours response to production incidents and critical service events, including evenings and weekends. Periodic travel may be required to support business needs and team on-sites.

Conditions of Employment:  Individual must successfully complete pre-employment process, which includes criminal background check, drug screening, credit check ( applicable for certain positions) and reference verification.

This job description reflects management’s assignment of essential functions. Nothing in this job description restricts management’s right to assign or reassign duties and responsibilities to this job at any time.

 

Caris Life Sciences is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, gender, gender identity, sexual orientation, age, status as a protected veteran, among other things, or status as a qualified individual with disability.

Skills

CI/CDDevOpsSREFDAPatient CareSOXCompliance

Similar Jobs

30

Software Engineer - Application Platform

Figma · San Francisco, CA • New York, NY • United States

1 week ago

Software Engineer – Application Platform

IMC · Sydney, Australia

1 week ago

Application Infra. Platform Engineer

Sandoz · Telangana (Sandoz), India

2 weeks ago

Software Engineer - Platform & Application

Bosonai · Santa Clara HQ · Onsite

1 month ago

Lead DevOps Engineer (Application & IT Operations)

Exadelinc · United States

Yesterday

Site Reliability Engineer (f/m/d) Application Hosting/TOSAAS

1&1 Drillisch · Revaler Straße 28-31, 10245 Berlin · Remote, Hybrid

5 days ago

Site Reliability Engineer (f/m/d) Application Hosting/TOSAAS

1&1 Drillisch · Hinterm Hauptbahnhof 3-5, 76137 Karlsruhe · Remote, Hybrid

5 days ago

Sr. Application Engineer (Platform Admin)

Employment Opportunities · Indianapolis, IN

1 week ago

Sr. Engineer Application and Development and Maintenance - ServiceNow Platform Technical Lead

Cardinalhealth · Philippines-Bonifacio Global City-Taguig

1 week ago

IT Infrastructure & Application Platform Support Engineer

Eurofins Scientific · Indaiatuba, SP, Brazil · Hybrid

2 weeks ago

Software Development Engineer II, Managed Application Platform

Amazon · Remote

2 weeks ago

AI Engineer: DevOps & Application Support

Haier · Appliance Park KY US, United States of America

2 weeks ago

Senior Application Support Engineer (SRE)

Depository Trust Company · Tampa, FL, United States, US

2 weeks ago

IT Infrastructure & Application Platform Support Engineer

Eurofins Scientific · Belo Horizonte, MG, Brazil · Hybrid

3 weeks ago

Sr. Application Engineer (DevOps, AI)

The Seattle Times · Seattle, WA · Remote, Onsite

3 weeks ago

Production Engineer, Site Reliability (Application Software)

Spacex · Hawthorne, CA +1

1 month ago

Software Engineer, Site Reliability Engineering (Application Software)

Spacex · Hawthorne, CA +1

1 month ago

Site Reliability Engineer (Application Software)

Spacex · Hawthorne, CA +1

1 month ago

Lead Platform Engineer - .NET, Desktop Application

Wells Fargo · 141278-NC-CIC Customer Information Ctr, United States of America

1 month ago

Application Engineer (Display & Electronics Product Platform)

3M · IN, Pune, India · Onsite

1 month ago

DevOps Engineer — KYC Application

Urban Connect · Bucharest · Hybrid

1 month ago

Technology Engineer Sr (DevOps/Linux, Web/Application servers, CI/CD )

PNC Bank · Two PNC Plaza (PA374), United States of America +1

1 month ago

Application Support DevOps Engineer

Recorded Future · Washington, DC +1

1 month ago

Application Support DevOps Engineer

Recorded Future · Boston, MA +1

1 month ago

Senior Site Reliability Engineer (Application)

Guidewire · Kuala Lumpur Office, Malaysia · Hybrid

1 month ago

Cloud Application Platform Delivery Engineer (Kubernetes)

Avaloq · Makati City, National Capital Region, Philippines · Hybrid

1 month ago

GIS Developer, DevOps, Geospatial Scripting & Application Support Opportunities

Xentity

1 month ago

Principal Application Support Engineer / SRE Lead (2nd Shift)

Depository Trust Company · Dallas, TX, United States, US

1 month ago

Senior Application Support Engineer / Site Reliability Engineer (SRE)

Depository Trust Company · Boston, MA, United States, US

1 month ago

Senior Application Security Platform Engineer

Rockstar Games · Manhattan, New York, United States · Onsite

1 month ago