Reliability & Observability Engineer (m/f/d) - #2730166
Certivity
Date: vor 3 Stunden
Stadt: München
Gehalt:
€75,000
-
€90,000
/ Jahr
Vertragstyp: Ganztags
Arbeitsplan: Volle Tag
About The Job
Our crawlers collect regulatory updates from 80+ sources and push them through an automated
pipeline: extraction, embedding, search, classification, consolidation and translation. As we add
regions, keeping it healthy has become a job of its own.
We want an engineer who makes production tell us what's wrong before customers do, fixes what
can be fixed automatically, and builds AI agents that diagnose the rest. When an issue does reach
a developer, it should arrive with a root cause and a suggested fix.
Not a ticket-driven ops role: you'll write production Python from week one and own platform
reliability alongside a core-team developer
What You Will Do
Python, Docker, GCP (Cloud Run, Cloud Logging, Cloud Monitoring), Sentry, MongoDB, Azure
Blob, LLM APIs. Some Go and React.
Your profile
OpenTelemetry, SLOs, Terraform or Pulumi, data-quality observability, crawling or document processing.
Why us?
We welcome candidates from all backgrounds and encourage diversity in our team. We encourage female and diverse engineers to apply and join our mission-driven culture that values open communication, work-life balance, and a welcoming environment. Convince us with your personality and your skills, and together we will make great things happen!
About Us
Certivity is a Munich-based RegTech startup using AI to turn global regulatory complexity into clear, structured, and actionable compliance intelligence. Our SaaS platform automatically gathers and analyzes regulatory data worldwide, enabling engineering and compliance teams to work faster and more confidently.
Our international team of 100+ experts supports over 15,000 users in 10 countries. While we are strongly established in the automotive sector, we are now expanding into new industries with high regulatory demands.
Our crawlers collect regulatory updates from 80+ sources and push them through an automated
pipeline: extraction, embedding, search, classification, consolidation and translation. As we add
regions, keeping it healthy has become a job of its own.
We want an engineer who makes production tell us what's wrong before customers do, fixes what
can be fixed automatically, and builds AI agents that diagnose the rest. When an issue does reach
a developer, it should arrive with a root cause and a suggested fix.
Not a ticket-driven ops role: you'll write production Python from week one and own platform
reliability alongside a core-team developer
What You Will Do
- Cut the noise. Separate transient failures (network blips, timeouts, errors that vanish on rerun) from real ones, and build alerting the team trusts.
- Build agentic incident response: agents that gather Sentry issues, logs, metrics and recent deploys, classify the failure, apply known fixes or open draft PRs, and brief the right
- Monitor the data, not just the infrastructure. A job that succeeds but extracts nothing is still a failure. Track freshness and completeness, e.g. "every source checked on time", "every
- Make the pipeline self-healing: retries with backoff, idempotent and resumable jobs, deadletter handling, clear escalation when automation gives up.
- Keep agents safe: scoped permissions, audit trails, human approval for risky actions, and
- Continuously audit our Infrastructure and Identify opportunities to make it more efficient and save costs.
- Catch memory, timeout and cost problems across Cloud Run before they become silent
- Add structured logging, metrics, tracing and sensible Sentry grouping, with infrastructure as code and CI/CD.
Python, Docker, GCP (Cloud Run, Cloud Logging, Cloud Monitoring), Sentry, MongoDB, Azure
Blob, LLM APIs. Some Go and React.
Your profile
- 5+ years in software engineering, SRE or platform roles, with real production Python.
- A track record of turning an ignored alert channel into one people act on.
- Hands-on experience building with LLMs or agents in production, and judgment about when
- Solid GCP (or similar), containers and CI/CD.
- Strong grasp of distributed-system failure modes: retries, idempotency, partial failure,
- Pragmatism, and clear communication in a small team.
OpenTelemetry, SLOs, Terraform or Pulumi, data-quality observability, crawling or document processing.
Why us?
- Hybrid work culture: join us in our Munich office (min. 2 days/week).
- Flexible working hours.
- 26+4 vacation days per year (4 fixed “company rest days” over Christmas).
- 30 days of “workation” per year, within the EU and selected countries.
- High autonomy and flat hierarchies.
- EGYM Wellpass for unlimited access to fitness courses and gyms.
- Udemy access for educational videos.
We welcome candidates from all backgrounds and encourage diversity in our team. We encourage female and diverse engineers to apply and join our mission-driven culture that values open communication, work-life balance, and a welcoming environment. Convince us with your personality and your skills, and together we will make great things happen!
About Us
Certivity is a Munich-based RegTech startup using AI to turn global regulatory complexity into clear, structured, and actionable compliance intelligence. Our SaaS platform automatically gathers and analyzes regulatory data worldwide, enabling engineering and compliance teams to work faster and more confidently.
Our international team of 100+ experts supports over 15,000 users in 10 countries. While we are strongly established in the automotive sector, we are now expanding into new industries with high regulatory demands.
Wie bewerbe ich mich?
Um sich für diesen Job zu bewerben, müssen Sie auf unserer Website autorisieren. Wenn Sie noch kein Konto haben, registrieren Sie sich bitte.
Veröffentlichen Sie einen LebenslaufÄhnliche Jobs
Regional Sales Leader, Marsh Risk Europe
Marsh,
vor 2 Stunden
Role purpose The Marsh Risk Regional Sales Leader, Europe is responsible for regional new business growth, pipeline outcomes, and sales execution across the region. The role provides commercial leadership to translate Marsh Risk growth priorities into clear sales actions, performance...
Sous Chef (m/w/d)
Hilton,
vor 2 Stunden
Job Description Exceptional Hospitality Starts with You Picture yourself brightening someone’s day. When you join our Hotels team, that’s exactly what you’ll do every time you come to work! As a Sous Chef , you’re not just leading daily kitchen...
Agentic Identity Sales Specialist
Palo Alto Networks,
vor 3 Stunden
Our Mission At Palo Alto Networks, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a...