Remote job
Principal Software Engineer, DevOps
Job details
About this role
Role overview A senior DevOps/SRE role responsible for the reliability, scalability, and cost-efficiency of a production cloud platform built on AWS and Kubernetes. The position blends infrastructure engineering with applied AI work, including building internal agentic systems that reduce operational toil and providing a safe platform for other engineers to build their own agents.
Responsibilities - Operate and evolve a public-cloud platform that spans AWS services, Kubernetes, containerization, and supporting infrastructure tooling - Build infrastructure-as-code and automation using Terraform, configuration management (e.g., Ansible/Chef), CI/CD pipelines (e.g., Jenkins/Git), and a programming language such as Go, Python, or Bash - Design, manage, and respond to production alerts; run root-cause investigations; and develop automated operational runbooks - Participate in an on-call rotation and support deployments during and outside regular business hours - Build and operate internal AI agents for incident triage, runbook execution, alert enrichment, and capacity/cost analysis, including their guardrails, rollback paths, and reliability metrics - Provide the platform that lets other engineers build agents safely: sandboxed execution, scoped credentials, tool/MCP integrations, cost controls, and agent observability
Requirements - 12+ years in SRE, DevOps, or SaaS operations - 3+ years managing AWS services and AWS cloud infrastructure with Terraform (AWS certification preferred) - 3+ years with containerized cloud solutions using Docker, Kubernetes, or AWS EKS, including familiarity with Helm charts - 5+ years with at least one programming language such as Go, Python, Bash, or Perl - 5+ years with infrastructure automation, CI/CD pipelines, and configuration management tools - Hands-on Linux/Unix experience (e.g., RedHat, CentOS, Ubuntu, Amazon Linux) and production experience building agentic AI systems, not just AI coding-assistant usage - Bachelor's or Master's degree in Computer Science/Engineering or a related field
Nice to have - Track record of introducing AI into an engineering SDLC with measurable outcomes, including judgment about where not to use it - 3+ years with observability stacks such as Splunk, Nagios, Elasticsearch, Kibana, CloudWatch, or Logstash, and experience scaling them - 3+ years with AWS RDS or Aurora MySQL - Demonstrated ability to collaborate with remote teams
Benefits and work setup - Base salary range $220,000–$255,000 plus stock equity, benchmarked using national and industry survey data and refined for candidate location and cost of living - Fully remote role based in the United States - Medical, dental, and vision plans with $0 monthly premium effective on day one - 401(k) plan, discretionary PTO and paid holidays, parental leave, equity, monthly wellness reimbursement, and a monthly lunch benefit