- Location
- Irving - HQ, United States of America
- Type
- Full-time
- Department
- Engineering
- Seniority
- Director
- Experience
- 10+ years
- Education
- Bachelor
- Source
- Workday
Description
At Caris, we understand that cancer is an ugly word—a word no one wants to hear, but one that connects us all. That’s why we’re not just transforming cancer care—we’re changing lives.
We introduced precision medicine to the world and built an industry around the idea that every patient deserves answers as unique as their DNA. Backed by cutting-edge molecular science and AI, we ask ourselves every day: “What would I do if this patient were my mom?” That question drives everything we do.
But our mission doesn’t stop with cancer. We're pushing the frontiers of medicine and leading a revolution in healthcare—driven by innovation, compassion, and purpose.
Join us in our mission to improve the human condition across multiple diseases. If you're passionate about meaningful work and want to be part of something bigger than yourself, Caris is where your impact begins.
Position Summary
Caris Life Sciences is one of the largest precision-oncology platforms in the world, serving hundreds of thousands of molecular cases a year and growing at double-digit rates. Behind every case is a matched molecular, imaging, and clinical-outcomes data estate few organizations anywhere can rival, and the clinical software that drives the lab's instruments and processes, captures results, and delivers each patient's report. When that software degrades, patient care waits; keeping it reliable is this role's charter.
Reporting to the Corporate Vice President for Clinical Software Products, the Director owns application reliability and production operations for the clinical software portfolio: production support, incident response, on-call operations, and SLO management for applications under SOX financial controls and FDA regulatory requirements. The Director hires and develops the team, sets standards and selects tooling, establishes production-access governance and segregation-of-duties controls with the information-security, quality, and infrastructure organizations, and participates directly in incident response. The role carries wide latitude to shape how reliability engineering is done here.
Frontier AI coding assistants are standard-issue tooling, with agentic workflows spanning incident diagnostics, runbook authoring, and operational automation. Characterization tests and golden-master replay validation serve as executable evidence for regulated change. Delivery runs on CI/CD with application-level observability and risk-based release governance aligned with FDA Computer Software Assurance guidance. Modernization of established systems is active engineering work, not deferred maintenance.
The infrastructure organization owns the platform and observability runtime; this role owns application-layer reliability on top of it. The operating model is automation-first: recurring manual work is engineered away rather than staffed, and operational and compliance evidence is produced by pipelines rather than assembled by hand.
Location: Irving, Texas (Dallas–Fort Worth), on-site/hybrid, minimum three days per week on campus; co-located with Caris's laboratory and clinical operations.
Job Responsibilities
Own the production-support model for the clinical application portfolio: the on-call rotation, escalation procedures, and incident-response playbooks.
Establish and audit production-access and segregation-of-duties controls with engineering, information-security, quality, and infrastructure partners; keep the evidence audit-ready for SOX ITGCs and applicable FDA requirements, including access grants, role changes, and privileged-action logs.
Define clinical SLOs, error budgets, and availability targets with product and engineering leadership; track attainment and operational-health metrics such as mean time to detect, mean time to recover, and on-call burden; intervene while an error budget is burning, not after it is spent.
Lead incident response for high-severity production events as incident commander or senior technical responder; coordinate cross-functional teams, run post-incident reviews, and drive systemic remediation to closure.
Develop the runbook library and approved operational automation; ensure engineers can execute standard interventions safely within documented procedures.
Drive the automation-first operating model: convert recurring manual interventions into reviewed automation, measure and reduce toil, and favor self-service tooling over ticket-driven request work.
Coordinate production deployments with engineering teams, verify deployment health, and own rollback decisions.
Advance automated generation of change and deployment evidence in CI/CD pipelines: deployment records, approval trails, and change documentation as an audit-ready by-product of release.
Hire and develop the App-SRE team; set performance expectations, on-call responsibilities, and career growth frameworks.
Own application-layer observability alongside the infrastructure and observability platform teams; keep production dashboards, alerting thresholds, and SLO monitors accurate and actionable.
Represent App-SRE in engineering leadership forums, operational reviews, and compliance audits; translate operational health and risk into clear executive communication.
Run the function AI-first: make AI-assisted practice the team's daily norm, from runbook automation to incident analysis and operational tooling.
Required Qualifications
Bachelor's degree in Computer Science, Software Engineering, Information Systems, or a closely related technical field, or equivalent practical experience.
10+ years of professional experience in SRE, DevOps, platform engineering, or production operations.
4+ years of direct experience in a people-management or team-lead role within an SRE or production-operations function.
Hands-on experience leading incident response for Tier 1 or business-critical production systems, including serving as incident commander or senior technical responder.
Experience defining and implementing SLOs, error budgets, and associated alerting and on-call workflows in a production environment.
Experience building or significantly maturing a production-support, on-call, or SRE function, including runbook development and on-call-rotation design.
Experience operating production systems under a formal regulatory or financial-controls framework, such as CAP/CLIA, FDA regulations, SOX ITGCs, HIPAA, or equivalent, including producing documentation that holds up in audit.
A record of applying AI-assisted practice to operations or engineering work, personally or through a team.
Preferred Qualifications
Domain experience in clinical diagnostics, laboratory information systems, molecular pathology, or digital health software.
Direct experience supporting SOX ITGC audit cycles or CAP/CLIA laboratory inspections, including evidence gathering for access-control, change-management, and monitoring controls.
Working knowledge of modern cloud-native observability at the application-instrumentation layer, including open standards for telemetry and tracing and application-performance-monitoring platforms.
Experience with deployment pipelines, release-management workflows, and rollback procedures in a continuous-delivery environment.
Ability to operate as a player-coach, contributing directly to technical work while building and leading a team.
Track record of reducing operational toil through automation programs in an SRE or production-operations organization.
Experience presenting operational strategy and risk posture to senior or executive audiences.
Physical Demands
Ability to sit, stand, and work at a computer for extended periods.
Training
All job-specific, safety, and compliance training is assigned based on the job functions associated with this employee.
Other
This role serves in the senior escalation tier of the production on-call rotation, with after-hours response to high-severity incidents as incident commander or senior escalation point. Periodic travel may be required to support business needs, team on-sites, and leadership reviews.
Conditions of Employment: Individual must successfully complete pre-employment process, which includes criminal background check, drug screening, credit check ( applicable for certain positions) and reference verification.
This job description reflects management’s assignment of essential functions. Nothing in this job description restricts management’s right to assign or reassign duties and responsibilities to this job at any time.
Caris Life Sciences is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, gender, gender identity, sexual orientation, age, status as a protected veteran, among other things, or status as a qualified individual with disability.