Remote job
Senior Site Reliability Engineer - US
Job details
About this role
Role overview A senior engineering role focused on building and operating a production SaaS platform that delivers secure access to customer infrastructure. The position centers on re-engineering core systems to scale globally, hardening reliability, and reducing operational toil. The work blends software engineering with traditional SRE practices in a security-first, 24/7 environment.
Responsibilities - Re-architect portions of the core product to enable global scaling and lower routing latency for distributed teams - Build out monitoring and observability tooling that surfaces real production issues while keeping alert noise to a minimum - Automate repetitive operational tasks to eliminate high-toil activities and free engineering capacity - Handle traditional operations work including patching, scaling, backup and restore, and disaster recovery - Investigate customer-facing outages and incidents, driving them to root cause and durable fixes - Participate in an on-call rotation supporting continuous system uptime
Requirements - 5+ years of progressive experience in software engineering, SRE, or DevOps roles - Strong background in Linux systems, networking, containers, and production troubleshooting - Solid hands-on experience programming in Go and working with Kubernetes - Proficiency with observability platforms such as Prometheus, Grafana, or Loki - Cloud experience, ideally with AWS, though GCP is acceptable - Strong communication skills, intellectual curiosity, and a collaborative, no-ego mindset; comfort working in a security-critical environment where correctness and system invariants matter
Nice to have - Experience writing automation scripts, contributing patches to a product codebase, or integrating AI agents into operational workflows - Familiarity with formal or property-based reasoning about system correctness
Benefits and work setup - Remote-first, globally distributed team with a one-week in-person onboarding in Oakland, CA - Extensive health coverage, an annual expense budget, retirement savings plans, and professional development support - Policies that prioritize rest and recovery, with autonomy to do meaningful work on a small team - Hiring process built around a take-home coding challenge in Go, a recruiter call, a hiring manager conversation, and a solution review