Remote job
Senior Site Reliability Engineer (Remote - India)
Job details
About this role
Role overview A senior engineering position focused on designing and operating highly available, scalable infrastructure across multiple cloud platforms. The role combines hands-on technical leadership with mentorship, driving reliability standards, automation strategy, and architectural decisions that shape platform-wide resilience and operational maturity.
Responsibilities - Architect and deliver reliable, cost-efficient infrastructure for mission-critical applications across multi-cloud environments - Define and refine SLOs, SLIs, and error budgets, establishing reliability standards adopted across teams - Build Infrastructure as Code solutions using Terraform, Ansible, and Azure Resource Manager templates or Bicep - Lead automation initiatives that reduce manual effort and enable self-service capabilities for development teams - Own incident response, conduct post-incident reviews, and drive systemic improvements to prevent recurrence - Mentor engineers, champion cloud-native patterns, and maintain documentation, runbooks, and knowledge repositories
Requirements - Significant senior-level experience operating production distributed systems at scale - Deep expertise with AWS and Azure cloud platforms - Strong proficiency with Infrastructure as Code tooling, particularly Terraform and Ansible - Demonstrated experience leading incident response and establishing SLO/SLI frameworks - Ability to influence technical direction, mentor engineers, and partner with cross-functional stakeholders - Solid grasp of networking, security, and observability fundamentals across complex environments
Nice to have - Familiarity with Kubernetes, service mesh technologies (Istio, Linkerd, Consul), and GitOps workflows - Exposure to chaos engineering tools such as Chaos Monkey or Gremlin - Understanding of disaster recovery, business continuity, and multi-region active-active architectures - Experience with MLOps, data pipeline orchestration, or supporting ML workloads in production - Knowledge of policy-as-code tools like Open Policy Agent and FinOps cost optimization practices