Hiring.Camp

Reliability & Observability Engineer (m/f/d)

Certivity

·

Yesterday

Location
München, Bayern
Type
Full-time
Department
Engineering
Source
Personio

Description

Company Culture


About the Job

Our crawlers collect regulatory updates from 80+ sources and push them through an automated 
pipeline: extraction, embedding, search, classification, consolidation and translation. As we add 
regions, keeping it healthy has become a job of its own.
We want an engineer who makes production tell us what's wrong before customers do, fixes what 
can be fixed automatically, and builds AI agents that diagnose the rest. When an issue does reach 
a developer, it should arrive with a root cause and a suggested fix.
Not a ticket-driven ops role: you'll write production Python from week one and own platform 
reliability alongside a core-team developer


Your mission


What you will do

• Cut the noise. Separate transient failures (network blips, timeouts, errors that vanish on rerun) from real ones, and build alerting the team trusts.
• Build agentic incident response: agents that gather Sentry issues, logs, metrics and recent deploys, classify the failure, apply known fixes or open draft PRs, and brief the right 
developer.
• Monitor the data, not just the infrastructure. A job that succeeds but extracts nothing is still a failure. Track freshness and completeness, e.g. "every source checked on time", "every 
document produced text and embeddings".
• Make the pipeline self-healing: retries with backoff, idempotent and resumable jobs, deadletter handling, clear escalation when automation gives up.
• Keep agents safe: scoped permissions, audit trails, human approval for risky actions, and 
measuring how often they're right.
• Continuously audit our Infrastructure and Identify opportunities to make it more efficient and save costs.
• Catch memory, timeout and cost problems across Cloud Run before they become silent 
OOM kills.
• Add structured logging, metrics, tracing and sensible Sentry grouping, with infrastructure as code and CI/CD.


Our Tech Stack

Python, Docker, GCP (Cloud Run, Cloud Logging, Cloud Monitoring), Sentry, MongoDB, Azure 
Blob, LLM APIs. Some Go and React.


Your profile

• 5+ years in software engineering, SRE or platform roles, with real production Python.
• A track record of turning an ignored alert channel into one people act on.
• Hands-on experience building with LLMs or agents in production, and judgment about when 
a plain if-statement is better.
• Solid GCP (or similar), containers and CI/CD.
• Strong grasp of distributed-system failure modes: retries, idempotency, partial failure, 
backpressure.
• Pragmatism, and clear communication in a small team.

Nice to have
OpenTelemetry, SLOs, Terraform or Pulumi, data-quality observability, crawling or document processing.


Why us?

  • Hybrid work culture: join us in our Munich office (min. 2 days/week).
  • Flexible working hours.
  • 26+4 vacation days per year (4 fixed “company rest days” over Christmas).
  • 30 days of “workation” per year, within the EU and selected countries.
  • High autonomy and flat hierarchies.
  • EGYM Wellpass for unlimited access to fitness courses and gyms.
  • Udemy access for educational videos.


Closing

We welcome candidates from all backgrounds and encourage diversity in our team. We encourage female and diverse engineers to apply and join our mission-driven culture that values open communication, work-life balance, and a welcoming environment. Convince us with your personality and your skills, and together we will make great things happen!

Skills

PythonReactAzureGCPDockerTerraformCI/CDMongoDBSRE

Similar Jobs

30

Site Reliability & Observability Engineer (SRE)

TransUnion·TUM - D.F. Del. Miguel Hidalgo, Mexico·Onsite

1d ago

Member of Technical Staff | Observability & Reliability

Avra·São Paulo·Remote

2w ago

Engineering Manager - Site Reliability & Observability (x/f/m)

Doctolib·Berlin, Germany +2

2mo ago

Engineering Manager - Site Reliability & Observability (x/f/m)

Doctolib·Paris, France +1

2mo ago

Engineering Manager - Observability & Reliability Engineering Obsession (x/f/m)

Doctolib·Berlin, Germany +1

8mo ago

Engineering Manager - Observability & Reliability Engineering Obsession (x/f/m)

Doctolib·Paris, France

11mo ago

Principal Observability & Reliability Architect

Ahead·US·Remote

11mo ago

Site Reliability Engineer - Observability Platform

Ford·US·Remote, Hybrid

1d ago

Site Reliability Engineer - Observability Platform

USA01 - USA Automotive·US·Remote

1d ago

Site Reliability Engineer - Observability

Adobe·Bucharest, Romania

3mo ago

Senior Software Engineer - Observability and Reliability

Sigma Computing·New York City, NY +1

3mo ago

Senior Software Engineer - Observability and Reliability

Sigma Computing·San Francisco, CA +1

3mo ago

Lead Site Reliability Engineer - Observability

SimCorp is·Hyderabad, India

4mo ago

Senior Site Reliability Engineer - Observability (x/f/m)

Doctolib·Berlin, Germany +1

4mo ago

DevOps Engineer – Observability & Platform Reliability

Amadeus·Lisbon - HQ, Portugal

5mo ago

Senior Site Reliability Engineer - Observability (x/f/m)

Doctolib·Berlin, Germany +2·Hybrid

1y+ ago

Site Reliability Engineer, AI Observability

Appnovation·New York, Austin +2

1d ago

Site Reliability Engineer, AI Observability

Appnovation·Toronto, Montreal +2

1d ago

Principal Site Reliability Engineer, Infrastructure Observability

Troweprice·London, Warwick Court

1mo ago

Reliability Engineer 4 (Observability Specialist )

Usbank·Chicago, IL +9

1mo ago

Observability Engineer - Site Reliability Engineering (SRE)

Ocbc·SGP-TC 2, Singapore·Onsite

1mo ago

Lead Observability and Site Reliability Engineer

Riyadhair·UNAVAILABLE

3mo ago

Principal Site Reliability Engineer, Infrastructure Observability

Troweprice·Owings Mills, MD - Building 3

6mo ago

Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)

Encora·Latin America

1w ago

Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)

Encora·Colombia

1w ago

Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)

Encora·Peru

1w ago

Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)

Encora·Mexico

1w ago

Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)

Encora·Brazil

1w ago

Senior Site Reliability Engineer (SRE) – OpenShift & Observability

Swift·OPC NL, Netherlands

1mo ago

Senior Site Reliability Engineer - Linux Systems & Application Observability

tastylive·Chicago, Illinois +1·Remote, Hybrid, Onsite

4w ago