Remote job
Site Reliability Engineer
Job details
About this role
Role overview
Help keep a customer-facing platform available, responsive, and recoverable when it operates directly on storefronts and other revenue-sensitive pages. The role spans production operations, deployment automation, monitoring, incident response, reliability targets, and practical security hardening, with hands-on responsibility for Linux-based infrastructure.
Responsibilities
- Operate deployment, monitoring, and alerting across production infrastructure, including VPS-based services and web process management. - Build CI/CD workflows that make routine releases dependable and rollbacks straightforward. - Define and measure meaningful availability and latency objectives tied to real user experience. - Lead incident response, communicate clearly during active disruptions, and produce postmortems that result in concrete fixes. - Strengthen TLS configuration, secrets handling, backup processes, and restore procedures. - Verify that backups and recovery plans work in practice rather than only existing as documentation.
Requirements
- At least four years in SRE, DevOps, or infrastructure engineering. - Strong Linux fundamentals and confidence operating production servers directly. - Experience designing or maintaining CI/CD pipelines and infrastructure as code. - Practical monitoring and observability experience, including actionable alerting. - Calm decision-making during incidents and clear written and verbal communication.
Nice to have
- Experience moving from VPS hosting to managed or containerized infrastructure. - Work optimizing infrastructure or model-inference costs. - Experience with security hardening or preparation for compliance requirements.