Remote job
Site Reliability Engineer
Job details
About this role
Role overview The Site Reliability Engineer owns the reliability, health, and deployment automation of a defense-oriented AI platform that supports mission-critical decision-making. This is a hands-on role focused on watching system telemetry, responding to incidents end-to-end, and maturing SRE practices so the platform can operate reliably at scale in secure environments.
Responsibilities - Monitor dashboards and telemetry to detect health issues, performance degradation, and reliability risks before they escalate into incidents. - Own the debugging and incident response process from first alert through resolution, deciding when to investigate further and when to escalate. - Build and improve logging, monitoring, and observability tooling that visualizes the state of the platform. - Develop and maintain platform health, scaling, and capacity planning practices. - Improve and automate build and deployment pipelines, and help build a secure, high-availability delivery pipeline. - Identify bottlenecks in engineering workflows and create self-service automation that makes the team faster and more reliable. - Communicate system status, trade-offs, and post-incident learnings clearly to teammates and stakeholders.
Requirements - 5+ years of experience in SRE, DevOps, or software engineering. - Hands-on experience monitoring production systems and running incident response, including on-call rotation and post-mortems. - Strong communication skills, especially the ability to stay clear and organized during live troubleshooting. - Experience with the Prometheus/Grafana/Loki stack, Datadog, or comparable enterprise observability tools. - Knowledge of AWS cloud technologies and infrastructure-as-code tools such as Terraform or Pulumi. - Proficiency in Python, Bash, or other scripting languages, plus comfort working in an agile scrum environment. - U.S. Citizenship and an active TS/SCI clearance; ability to work on-site in San Diego, CA.
Nice to have - Defining and tracking SLOs, SLIs, and error budgets. - Crafting CI/CD pipelines and automation, and using Docker or other containerization tooling. - Working in AWS GovCloud, modern web service architectures, relational databases, or Elasticsearch/OpenSearch.
Benefits and work setup - Health, dental, and vision insurance. - 100% remote-first culture with flexible work from anywhere in the U.S., plus shared workspace access. - Unlimited paid time off and competitive holiday schedules. - Monthly wellness, mental health, and home-office stipends, plus family planning assistance. - Salary top-up during military reserve duty, fully paid parental leave, and child/pet care reimbursement during travel.