Remote job
Senior Site Reliability Engineer
Job details
About this role
Role overview
A cloud application platform company is hiring a Senior Site Reliability Engineer to evolve its platform from traditional cloud operations into a proactive, automation-driven SRE model across multi-cloud environments. The role owns critical engineering workstreams that improve reliability, scalability, and operational efficiency, and partners with engineering, product, and platform teams to embed reliability into the software delivery lifecycle. The team is remote and globally distributed, with follow-the-sun support and a B Corp ethos.
Responsibilities
- Architect system monitoring, alerting, and logging using Prometheus, Grafana, and the ELK Stack, and define SLIs/SLOs tied to business metrics.
- Reduce operational toil by designing resilient, automated infrastructure with Terraform and Ansible across AWS, GCP, and Azure.
- Optimize CI/CD pipelines for fast, secure, zero-downtime releases under high-volume deployment cycles.
- Lead high-priority incident triage, drive blameless post-mortems, and implement preventative measures.
- Partner with product and software engineering to bring SRE best practices into product roadmaps.
- Identify performance bottlenecks and evaluate emerging technologies such as eBPF and container orchestration.
Requirements
- 5+ years in Site Reliability Engineering, Cloud Operations, or DevOps, with proven ownership of reliability for production platforms at scale.
- Strong proficiency in Go or Python to build automation, custom controllers, or SRE platform components beyond shell scripting.
- Advanced hands-on knowledge of Linux internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting.
- Deep expertise with cloud providers (AWS, GCP, Azure, or OpenStack), cloud SDKs, and declarative infrastructure tools such as Terraform.
- Proven ability to anticipate operational risks, make architectural trade-offs, and lead infrastructure initiatives with minimal guidance.
- Strong cross-functional communication and a track record of building alignment and an inclusive engineering culture.
Nice to have
- Experience with custom-built orchestration, edge, storage, or operational tooling in a dynamic environment.
Benefits and work setup
- Fully remote role with flexible PTO, company stock options, a professional development budget, office equipment budget, wellness budget, internet reimbursement, inclusive parental leave, an annual team gathering, and a remote work travel program.
- On-call rotation roughly every 4–5 weeks, with shifts including 02:00–10:00 UTC and a weekend component during the on-call week.
- Hiring process includes a talent acquisition screen, hiring manager interview, team interview, and senior director interview; background checks are required.