Remote job
Site Reliability Engineer (f/m/d) – Observability & Internal Tools
Job details
About this role
Role overview A platform-focused Site Reliability Engineer position dedicated to elevating observability and internal tooling into a true engineering capability. The engineer acts as guardian and architect of internal infrastructure, designing systems rather than consuming off-the-shelf SaaS, with full end-to-end ownership of the stack. Work is remote by default, with periodic on-site collaboration for architectural decisions and larger initiatives.
Responsibilities - Own and evolve the observability ecosystem centered on Prometheus, Grafana, and Forgejo, turning monitoring tools into a coherent platform-grade capability. - Design actionable, signal-rich alerts and define Service Level Objectives that shift operations from reactive firefighting toward proactive stability. - Evaluate and integrate modern open-source alternatives to proprietary software, influencing technology selection across the stack. - Embed security engineering directly into the development and delivery pipeline, identifying vulnerabilities earlier than external penetration testing. - Operate deep within Linux systems and distributed tooling, balancing bold experimentation with a strictly stable production environment. - Enable and accelerate the productivity of other engineers through automation and platform improvements.
Requirements - Demonstrable hands-on experience operating Prometheus, Grafana, and related observability tooling in production environments. - Solid grasp of Linux internals, distributed systems, and infrastructure-as-code practices. - Track record of building rather than merely configuring, with a portfolio, side project, or production-ready demo repository that shows shipping ability. - Comfort defining SLOs, crafting meaningful alerts, and moving an organization from reactive to proactive reliability work. - Practical knowledge of integrating security practices into CI/CD and delivery workflows. - Builder mindset with high standards, low ego, openness to direct feedback, and the ability to move fast while staying careful where it counts.
Nice to have - Active engagement with the open-source community, hackathons, or conference participation.
Benefits and work setup - Remote-first working model with on-site collaboration in Germany for key planning and brainstorming sessions. - 30 days of vacation plus time off on Christmas Eve and New Year's Eve. - Optional four-day work week through dedicated focus Fridays. - Mobility benefits including a public-transport subsidy and a bicycle leasing program. - Sports, health, and mental-health support offerings. - Corporate benefits platform with retail and lifestyle discounts. - Investment in continued learning through community events, hackathons, and conferences.