← Back to jobs

Remote job

Senior Site Reliability Engineer

DevOps Full-time Permanent US, AR, BR, CA

Job details

$152,000 - $195,000 Salary
US, AR, BR, CA Eligibility
Senior Experience
Full-time Employment

About this role

Role overview A senior site reliability engineering role centered on Kubernetes-based infrastructure, CI/CD pipelines, and the operational foundation behind AI tooling for production systems. The position combines hands-on platform engineering with technical leadership, including owning AI access patterns, observability, and resilience across multi-tenant environments.

Responsibilities - Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications - Build AI tooling infrastructure, including standing up MCP servers and defining safe, auditable access patterns for production systems - Optimize and maintain CI/CD pipelines for reliability, speed, and rollback safety, including progressive delivery strategies such as blue/green and canary deployments - Advance Infrastructure as Code practices using Terraform, Helm, and Argo CD, defining reusable patterns across the organization - Operate and optimize streaming and analytics infrastructure spanning Kafka, Flink, and ClickHouse - Improve system observability through SLOs, alerts, dashboards, and postmortem-driven root cause analysis, while mentoring engineers on Kubernetes, CI/CD, and cloud infrastructure

Requirements - 6+ years in SRE, DevOps, or infrastructure roles with significant production Kubernetes experience, including managed services such as EKS, GKE, or AKS - Hands-on experience integrating AI or LLM tooling into engineering or operational workflows, with a clear grasp of the security and governance considerations of giving AI access to production - Proven track record building CI/CD pipelines with tools such as GitHub Actions, Jenkins, or GitLab CI - Strong Infrastructure as Code expertise with Terraform, Helm, or Pulumi and GitOps practices - Proficiency in Python, Bash, or Go - Knowledge of observability tooling such as Prometheus, Grafana, Datadog, or OpenTelemetry, plus production experience with Kafka, Flink, and ClickHouse

Nice to have - Multi-region or multi-cluster Kubernetes experience - Chaos engineering or resilience testing background - Security scanning, compliance automation, or policy-as-code experience - Familiarity with LLM observability and tracing tooling such as LangSmith or Langfuse, or with MLOps workflows - Contributions to open-source Kubernetes or CI/CD projects

Benefits and work setup - Estimated total compensation range of $152,000 to $195,000, including base plus bonus, with eligibility for annual performance-based incentive awards and equity - Competitive salary, stock options, health benefits, unlimited paid time off, parental leave, and tuition reimbursements - Commitment to equal employment opportunity and reasonable accommodations for candidates with disabilities - Note: immigration sponsorship is not provided for this position

Skills detected in the listing

PythonGoInformation SecurityKubernetesTerraformLLM
Detected Sep 16, 2026
Last verified Not yet verified

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight