← Back to jobs

Remote job

Staff Site Reliability Engineer

DevOps Full-time Permanent US (Remote)

Job details

Not specified Salary
US (Remote) Eligibility
Staff Experience
Full-time Employment

About this role

Role overview

This staff-level SRE position is foundational to setting engineering-wide reliability strategy and standards at a remote cybersecurity company. The role owns the SRE operating model, drives cross-functional reliability initiatives, and shapes technical direction across infrastructure, product, service, and security teams supporting a production-safe autonomous pentest platform.

Responsibilities

- Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards in alignment with customer impact and business priorities - Lead cross-functional alignment across infrastructure, product, security, and business stakeholders to improve reliability, observability, incident response, and operational readiness - Establish organization-wide approaches to service ownership, SLIs, SLOs, and error budgets for critical customer paths - Define observability standards and set the bar for dashboards, actionable alerting, runbooks, and escalation paths - Drive end-to-end complex reliability initiatives spanning multiple teams - Raise the engineering-wide standard for incident management, incident command, on-call health, and post-incident learning - Shape the technical direction, operating model, and growth path of the SRE function - Participate in a 24/7 on-call rotation and help design a sustainable, appropriately staffed on-call model

Requirements

- Experience designing, operating, and troubleshooting large-scale distributed systems in production - Deep knowledge of reliability engineering, observability, incident management, and production operations, with a track record of turning that knowledge into adopted standards - Demonstrated experience establishing SLIs, SLOs, actionable alerts, observability, and service ownership models - Backend experience building automation that reduces toil, strengthens safeguards, and improves operational efficiency - Experience leading high-severity incidents and improving incident response programs - Strong written and verbal communication skills for technical designs, runbooks, postmortems, and operational documentation - Proficiency with Python and Terraform, or equivalent infrastructure-as-code and automation tooling - Hands-on experience with observability platforms such as Datadog, New Relic, or Grafana - Production experience operating services on AWS and Kubernetes - Familiarity with CI/CD pipelines such as GitLab CI, ArgoCD, or GitOps workflows

Benefits and work setup

- Base salary range of $199,750 to $270,000 annually, depending on location, qualifications, experience, and relevant skills - Eligibility for an equity package in the form of stock options for all full-time roles - Fully remote with up to 10% travel for team off-sites and in-person project kick-offs - Comprehensive benefits including medical and dental insurance for the employee and family, flexible vacation policy, and generous parental leave

Skills detected in the listing

PythonStakeholder ManagementInformation SecurityAWSKubernetesTerraform
Detected Sep 25, 2026
Last verified Sep 25, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight