Remote job
SRE, DevOps, Cloud, Platform Engineering, CI/CD
Job details
About this role
Role overview
This remote Site Reliability Engineering role centers on evolving and optimizing a production platform that supports cloud infrastructure management, deployment automation, and DevOps workflows. You will help make systems more reliable, efficient, observable, and cost-conscious while reducing repetitive operational work through automation. The role includes meaningful influence over production infrastructure, platform architecture, developer experience, and reliability practices.
Responsibilities
- Maintain a reliability-first approach to production systems and operational support. - Diagnose and resolve issues affecting production applications and environments. - Improve infrastructure provisioning automation and overall service availability. - Identify opportunities to increase performance and efficiency while reducing infrastructure costs. - Automate repetitive operational tasks and reduce manual toil for engineering teams. - Work with open-source and private repositories and integrate with cloud application and infrastructure services.
Requirements
- Experience with SRE, DevOps, cloud infrastructure, platform engineering, or CI/CD practices. - Ability to debug production applications and infrastructure environments. - Understanding of reliability, availability, observability, automation, and operational efficiency. - Proactive approach to making architectural and production-infrastructure decisions. - Strong interest in developer experience and simplifying engineering workflows.
Nice to have
- Experience with major cloud providers, Docker, Kubernetes, REST APIs, shell scripting, or infrastructure automation. - Familiarity with distributed application environments, monitoring, databases, or deployment systems. - Experience working across a broad technology ecosystem, including services built with Go, Python, JavaScript, TypeScript, or similar languages.
Benefits and work setup
- Remote position with full-time or part-time availability. - Collaborative work with opportunities to address operational problems as shared learning and improvement initiatives.