Remote job
Senior Site Reliability Engineer
Job details
About this role
Role overview A senior site reliability engineering role centered on Kubernetes-based infrastructure, CI/CD pipelines, and the operational foundation behind AI tooling for production systems. The position combines hands-on platform engineering with technical leadership, including owning AI access patterns, observability, and resilience across multi-tenant environments.
Responsibilities - Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications - Build AI tooling infrastructure, including standing up MCP servers and defining safe, auditable access patterns for production systems - Optimize and maintain CI/CD pipelines for reliability, speed, and rollback safety, including progressive delivery strategies such as blue/green and canary deployments - Advance Infrastructure as Code practices using Terraform, Helm, and Argo CD, defining reusable patterns across the organization - Operate and optimize streaming and analytics infrastructure spanning Kafka, Flink, and ClickHouse - Improve system observability through SLOs, alerts, dashboards, and postmortem-driven root cause analysis, while mentoring engineers on Kubernetes, CI/CD, and cloud infrastructure
Requirements - 6+ years in SRE, DevOps, or infrastructure roles with significant production Kubernetes experience, including managed services such as EKS, GKE, or AKS - Hands-on experience integrating AI or LLM tooling into engineering or operational workflows, with a clear grasp of the security and governance considerations of giving AI access to production - Proven track record building CI/CD pipelines with tools such as GitHub Actions, Jenkins, or GitLab CI - Strong Infrastructure as Code expertise with Terraform, Helm, or Pulumi and GitOps practices - Proficiency in Python, Bash, or Go - Knowledge of observability tooling such as Prometheus, Grafana, Datadog, or OpenTelemetry, plus production experience with Kafka, Flink, and ClickHouse
Nice to have - Multi-region or multi-cluster Kubernetes experience - Chaos engineering or resilience testing background - Security scanning, compliance automation, or policy-as-code experience - Familiarity with LLM observability and tracing tooling such as LangSmith or Langfuse, or with MLOps workflows - Contributions to open-source Kubernetes or CI/CD projects
Benefits and work setup - Estimated total compensation range of $152,000 to $195,000, including base plus bonus, with eligibility for annual performance-based incentive awards and equity - Competitive salary, stock options, health benefits, unlimited paid time off, parental leave, and tuition reimbursements - Commitment to equal employment opportunity and reasonable accommodations for candidates with disabilities - Note: immigration sponsorship is not provided for this position