Remote job
Staff Site Reliability Engineer - US
Job details
About this role
Role overview A senior engineering position focused on building the production infrastructure for a software-as-a-service security and access platform from the ground up. The role centers on re-engineering core product components to scale globally, reduce routing latency for distributed users, and support a cloud delivery model. Most code is written in Go, and security rigor is non-negotiable because a breach could put customer infrastructure at risk.
Responsibilities - Re-engineer core product components to scale globally and minimize routing latency for geographically distributed teams - Rewrite portions of the core product to enable cloud product goals and multi-region delivery - Build and refine the monitoring and observability stack so the team detects real issues without alert fatigue - Automate the highest-toil operational workflows to reduce manual intervention - Handle traditional operations work including patching, scaling, backup and restore, and disaster recovery - Investigate customer-facing outages and incidents, and participate in a 24/7/365 on-call rotation
Requirements - 8 or more years of progressive experience in software engineering, SRE, or DevOps roles - Strong background in Linux systems, networking, containers, and production troubleshooting - Solid Go programming skills along with production Kubernetes development experience - Demonstrated technical leadership, including mentoring, design review, or owning critical systems - Hands-on experience operating and supporting an observability platform using tools such as Prometheus, Grafana, or Loki - Strong scripting and automation skills, ideally with experience submitting patches to a product codebase or building tooling that incorporates AI agents into operational workflows - Track record of working in security-sensitive environments where reasoning about correctness and system invariants is valued - Excellent communication skills and intellectual curiosity, with a transparent, no-ego working style - Willingness to complete a collaborative take-home coding challenge in Go as part of the interview process
Benefits and work setup - Remote-first team with globally distributed colleagues - Required in-person attendance for one onboarding week in Oakland, California - Extensive health coverage, an annual expense budget, and rest-and-recovery policies designed to prevent burnout - Retirement savings plan and ongoing professional development support - Background checks conducted as part of the hiring process