Remote job
SRE, DevOps, Cloud, Platform Engineering, CI/CD
Job details
About this role
Role overview
Join a remote engineering team responsible for improving the reliability, performance, efficiency, and operational maturity of a production platform used for cloud infrastructure and deployment automation. This Site Reliability Engineering role combines production debugging, infrastructure automation, observability, cost awareness, and the reduction of repetitive operational work. You will influence production architecture and help the team scale dependable services for users around the world.
Responsibilities
- Keep production systems reliable and respond effectively to application and environment issues. - Diagnose and debug production applications, infrastructure, and deployment problems. - Improve infrastructure provisioning automation, availability, performance, and operational efficiency. - Reduce manual and repetitive work by automating recurring operational tasks and workflows. - Contribute to reliability, observability, cost-reduction, and platform-engineering initiatives. - Work with internal and third-party cloud infrastructure services across open-source and proprietary repositories.
Requirements
- Practical experience in SRE, DevOps, cloud infrastructure, platform engineering, or CI/CD. - Strong troubleshooting skills and confidence investigating production systems. - Interest in automation, reliability engineering, observability, infrastructure management, and developer experience. - Ability to influence architecture and make sound decisions about production infrastructure. - Proactive, ownership-oriented approach to operational problems and continuous improvement.
Nice to have
- Experience with cloud providers, Docker, Kubernetes, infrastructure provisioning, or shell scripting. - Familiarity with application stacks, databases, caching systems, and REST APIs.