Remote job
Sr Site Reliability Engineer
Job details
About this role
Role overview A senior engineering role focused on cloud platform reliability, joining a Platform Infrastructure team responsible for a microservices-based software solution built on modern orchestration technologies. The position emphasizes leading the engineering organization toward standardized, automated infrastructure and service provisioning, while operating within a culture that values flexibility, trust, and continuous learning. The work blends hands-on engineering with cross-team technical leadership and on-call incident response.
Responsibilities - Architect long-term technical solutions and cross-team mechanisms that move the platform toward measurable reliability goals. - Define and drive a roadmap for self-service, automated, scalable, and observable infrastructure services delivered as a product. - Align stakeholders and lead execution of the Platform Infrastructure team's roadmap alongside senior engineers across the organization. - Provide expert guidance and review during engineering design sessions for teams onboarding to the shared platform. - Aggressively reduce manual toil through automation across provisioning, deployment, and operations. - Build and maintain monitoring and alerting for the platform, and participate in an on-call rotation for incident response.
Requirements - Substantial experience designing technical solutions for reliability at scale in a cloud-native, microservices environment. - Demonstrated ability to lead roadmap-level initiatives and influence engineering practices across multiple teams. - Hands-on expertise with AWS services such as S3, EC2, and RDS, including EKS-managed Kubernetes clusters. - Practical knowledge of service mesh technologies (Istio) and Infrastructure as Code tools such as Terraform or AWS CDK. - Familiarity with continuous integration in GitLab and continuous delivery pipelines using tooling like ArgoCD. - Experience with observability platforms such as Datadog for monitoring and alerting.
Benefits and work setup - Culture built around flexibility, trust, and continual learning, with explicit emphasis on diversity and inclusion as guiding values. - On-call rotation is part of the role, suggesting structured incident response and operational ownership expectations. - Position framed as a leadership-track opportunity within the platform infrastructure function.