Hiring.Camp

Reliability & Incident Manager (m/f/d)

CGM Career

·

Apr 29, 2026

Location
Koblenz | Maria Trost 21, Germany
Type
Full-time
Seniority
Manager
Source
Workday

Description

As a leading provider of software solutions for healthcare, we operate in 19 countries and employ nearly 9,000 dedicated professionals. You will work in a dynamic and innovative environment, filled with exciting opportunities. With your commitment and passion, you’ll have the chance to make a lasting impact.

 

CGM Leverages AI: We are looking for people who are inspired by the power of AI in eHealth, eager to shape transformation, and curious at heart - ready to see how technology can make healthcare smarter, easier, and better.

 

Together, we are shaping the future of healthcare. Become part of our mission and make a difference - for a world where knowledge saves lives!

 

In this role, you are responsible for the end-to-end experience of our customers in support and service, from the first contact to the final resolution. You don’t just look at individual touchpoints, but at the entire journey: intake, routing, processing, resolution, and feedback. The ambition behind this is clear: our customers should not experience internal complexity – they should experience service. That’s why you systematically analyze friction points, drive improvements based on data, and deliberately leverage tools, standards, and automation to measurably increase customer satisfaction.

Your contribution:

  • Analyze recurring incidents, support drivers, and release defects with a strong focus on identifying and eliminating root causes rather than symptoms, directly contributing to a continuous reduction in incident recurrence by at least 1 percentage point per month
  • Lead and facilitate the RQIL Review Board as the central operational mechanism for incident prioritization, root cause validation, and enforcement of corrective actions, ensuring measurable cycle times below 14 days from identification to prevention closure
  • Ensure that every incident resolution is translated into durable prevention artifacts such as test coverage, release gates, telemetry improvements, runbooks, or engineering patterns, directly reducing Change Failure Rate to below 5%
  • Partner closely with Product and Engineering to define hardening initiatives and support data driven trade offs between reliability, delivery, and roadmap priorities
  • Identify and close telemetry gaps, improving observability and enabling earlier detection of system issues before they impact customers, increasing overall telemetry coverage to at least 80%
  • Drive structured postmortems for critical incidents, including clear ownership, defined actions, and measurable verification of effectiveness
  • Develop and maintain preventive playbooks for recurring risk patterns across products and platforms, ensuring systemic issues are removed rather than repeatedly mitigated

What you bring:

  • Several years of experience in Reliability Engineering, Incident Management, DevOps, or similar technical operations roles in complex software environments
  • Strong root cause mindset with the ability to consistently move from symptoms to systemic issues
  • Proven experience in establishing and operating structured incident, postmortem, or SRE frameworks
  • Solid understanding of observability, telemetry, correlation IDs, and modern monitoring architectures
  • Ability to operate effectively with Engineering and Product teams, influencing technical priorities through data and structured reasoning
  • Experience in defining and enforcing operational standards, processes, and governance models across teams or products
  • Familiarity with ITSM and delivery tools such as ServiceNow, Jira, or comparable platforms

What you can expect from us:

  • Mobile Work: Work flexibly from home two days a week and on-site three days a week

  • Attractive locations: Our offices offer fully equipped workspaces as well as regular events such as summer parties and Christmas celebrations

  • Development: Our in-house academy and a portfolio of external partners support your professional growth

  • Health: We place high value on health. At our in-house canteen in Koblenz, you’ll find a daily selection of delicious and healthy meals, and our fully equipped gym offers weekly classes (online & on-site)

  • And more: A daycare center on our CGM campus in Koblenz helps employees make their workday more flexible. We also offer corporate benefits, the option of a job bike, company pension schemes, and much more

Diversity is part of CGM! We welcome applications regardless of disability, gender, nationality, ethnic and social background, religion, age, sexual orientation, or identity.

We are looking for people who recognize the power of AI in the eHealth environment, want to help shape change, and are driven by a curious passion to understand how technology can make healthcare smarter, simpler, and better.

Interested? Apply now online with your meaningful application documents (including all certificates, salary expectations, and your earliest possible starting date).

Skills

JiraServiceNowDevOpsSRE

Similar Jobs

11

Incident Management Reliability Engineer

Sanofi · Budapest, Hungary · Remote, Hybrid

3 weeks ago

Senior Incident Optimization & Reliability Specialist - End-User Technology – Vice President

Citi Bank · TRIL INFO PARK, LITTLEWOOD TOWER, India · Hybrid

1 month ago

Senior Incident Optimization & Reliability Specialist - End-User Technology – Vice President

citibank · Chennai, TN,IN, IN

1 month ago

Incident management / reliability / SRE Evaluator

Mercor · US · Remote

1 month ago

Site Reliability Engineer (Incident Manager)

Arkoselabs · Brisbane, Australia +1 · Remote, Hybrid

1 month ago

Associate Site Reliability Engineer (Incident Manager)

Arkoselabs · Brisbane, Australia +1 · Remote, Hybrid

1 month ago

Staff Site Reliability Engineer - Incident Management & Reliability (Remote - Canada)

Confluent · Remote, Ontario, Canada +1 · Remote

5 months ago

Site Reliability Engineer II - Incident Management

JPMorgan Chase · Hyderabad, Telangana, India

4 days ago

Site Reliability Engineer II - Incident Management

JP Morgan Chase · Hyderabad, Telangana, India

4 days ago

Site Reliability Engineer II – AWS, Incident Response, Automation, Observability

JPMorgan Chase · Hyderabad, Telangana, India

2 months ago

Senior Incident & Automation Engineer (AIOps / Reliability) Vice President

Citi Bank · 6400 LAS COLINAS BLVD IRVING, United States of America

2 months ago