Remote job
Senior Site Reliability Engineer
Job details
About this role
Role overview A senior engineering role focused on evolving a multi-cloud application platform toward a proactive, automation-first reliability model. The position partners closely with engineering, product, and platform teams to embed reliability and performance into every stage of the software delivery lifecycle for a global user base. Work spans infrastructure-as-code, observability, incident leadership, and architectural improvements across AWS, GCP, and Azure.
Responsibilities - Design and elevate monitoring, alerting, and logging across distributed systems using Prometheus, Grafana, and the ELK Stack, defining SLIs and SLOs tied to core business outcomes. - Reduce operational toil by building resilient automation with Terraform and Ansible, and by extending platform tooling through custom Go or Python components. - Optimize CI/CD pipeline architecture to deliver fast, secure, zero-downtime releases that hold up under heavy deployment volume. - Lead high-priority incident triage, drive blameless post-mortems, and convert findings into preventative reliability improvements. - Partner with product and software engineering teams to shape roadmaps around SRE best practices, performance budgets, and capacity planning. - Evaluate emerging technologies such as eBPF and advanced container orchestration to address bottlenecks before they affect customers.
Requirements - At least 5 years in Site Reliability Engineering, Cloud Operations, or DevOps, with a track record of owning reliability for production platforms at scale. - Strong proficiency in Go or Python for building custom automation, controllers, or platform components beyond shell scripting. - Deep hands-on knowledge of Linux internals, kernel parameters, networking protocols, performance profiling, and system-level troubleshooting. - Deep experience with AWS, GCP, Azure, or OpenStack, including declarative infrastructure tooling such as Terraform to manage distributed systems. - Proven ability to anticipate operational risks, weigh architectural trade-offs, and drive technical infrastructure initiatives independently. - Strong cross-functional communication skills with a history of building alignment and fostering an inclusive engineering culture.
Nice to have - Experience building custom orchestration, edge, storage, or operational tooling in a fast-moving environment.
Benefits and work setup - Fully remote position within a globally distributed team. - Flexible paid time off plus company stock options. - Annual budgets for professional development, office equipment, and wellness. - Internet reimbursement, inclusive parental leave, and a remote work travel program. - Periodic in-person team gatherings. - On-call rotation roughly one week every 4–5 weeks, including a weekend shift, with the on-call window running 02:00–10:00 UTC. - Candidates must be legally authorized to work in the listed countries; visa sponsorship is not currently offered. - Hiring process includes four Google Meet interviews followed by a background check.