Remote job
Staff Site Reliability Engineer
Job details
About this role
Role overview
This staff-level SRE position is foundational to setting engineering-wide reliability strategy and standards at a remote cybersecurity company. The role owns the SRE operating model, drives cross-functional reliability initiatives, and shapes technical direction across infrastructure, product, service, and security teams supporting a production-safe autonomous pentest platform.
Responsibilities
- Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards in alignment with customer impact and business priorities - Lead cross-functional alignment across infrastructure, product, security, and business stakeholders to improve reliability, observability, incident response, and operational readiness - Establish organization-wide approaches to service ownership, SLIs, SLOs, and error budgets for critical customer paths - Define observability standards and set the bar for dashboards, actionable alerting, runbooks, and escalation paths - Drive end-to-end complex reliability initiatives spanning multiple teams - Raise the engineering-wide standard for incident management, incident command, on-call health, and post-incident learning - Shape the technical direction, operating model, and growth path of the SRE function - Participate in a 24/7 on-call rotation and help design a sustainable, appropriately staffed on-call model
Requirements
- Experience designing, operating, and troubleshooting large-scale distributed systems in production - Deep knowledge of reliability engineering, observability, incident management, and production operations, with a track record of turning that knowledge into adopted standards - Demonstrated experience establishing SLIs, SLOs, actionable alerts, observability, and service ownership models - Backend experience building automation that reduces toil, strengthens safeguards, and improves operational efficiency - Experience leading high-severity incidents and improving incident response programs - Strong written and verbal communication skills for technical designs, runbooks, postmortems, and operational documentation - Proficiency with Python and Terraform, or equivalent infrastructure-as-code and automation tooling - Hands-on experience with observability platforms such as Datadog, New Relic, or Grafana - Production experience operating services on AWS and Kubernetes - Familiarity with CI/CD pipelines such as GitLab CI, ArgoCD, or GitOps workflows
Benefits and work setup
- Base salary range of $199,750 to $270,000 annually, depending on location, qualifications, experience, and relevant skills - Eligibility for an equity package in the form of stock options for all full-time roles - Fully remote with up to 10% travel for team off-sites and in-person project kick-offs - Comprehensive benefits including medical and dental insurance for the employee and family, flexible vacation policy, and generous parental leave