Remote job
Platform Reliability Engineer (m/w/d) - Konstanz, Berlin oder remote
Job details
About this role
Role overview A Platform Reliability Engineer role focused on building and operating resilient infrastructure that supports software development services. The position combines platform engineering with site reliability practices, working on Infrastructure-as-Code and Kubernetes-based environments. It is a permanent, full-time opportunity based in Konstanz, Berlin, or fully remote.
Responsibilities - Design, build, and maintain platform infrastructure using Kubernetes and Infrastructure-as-Code tooling such as Pulumi. - Improve system reliability, availability, and performance across the platform layer. - Respond to and manage incidents, driving root cause analysis and preventive follow-ups. - Contribute to software and system architecture decisions that shape how services are deployed and operated. - Develop and evolve infrastructure components and shared platform capabilities for engineering teams. - Document operational practices and support the broader engineering organization in adopting the platform.
Requirements - Proven experience in a platform, reliability, or infrastructure engineering role. - Strong hands-on knowledge of Kubernetes in production environments. - Practical experience with Infrastructure-as-Code, ideally including Pulumi, and related patterns for managing cloud resources. - Familiarity with incident management processes and reliability engineering practices. - Comfort working within software and system architecture contexts, bridging development and operations. - Ability to work independently in a remote or hybrid setup, with professional-level communication skills.
Nice to have - Background in a software development and services organization, working closely with product engineering teams. - Experience shaping platform offerings consumed by multiple internal stakeholders.