Remote job
Site Reliability Engineer
Job details
About this role
Role overview
The SRE will own reliability, availability, and performance of production systems running in cloud environments that power a workflow orchestration platform handling billions of mission-critical transactions. The role blends incident response, observability, automation, and cross-team partnership to scale distributed systems used across fintech, e-commerce, logistics, and healthcare.
Responsibilities
- Own reliability, availability, and performance of production cloud systems. - Define and monitor SLIs and SLOs and manage error budgets across the platform. - Lead incident response including detection, triage, mitigation, and postmortems. - Improve observability through logging, monitoring, alerting, and dashboards. - Automate operational workflows and reduce manual toil wherever possible. - Partner closely with engineering teams to improve system resiliency and scalability. - Support capacity planning, infrastructure optimization, and performance tuning. - Build internal tooling, runbooks, and operational best practices. - Support Kubernetes-based infrastructure and distributed systems at scale. - Act as escalation point for complex production and platform issues.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure roles. - Strong experience with cloud platforms such as AWS, GCP, or Azure. - Hands-on experience with Kubernetes and containerized environments. - Strong understanding of distributed systems and microservices architecture. - Experience with observability tools such as Prometheus, Grafana, Datadog, ELK, or OpenTelemetry. - Proficiency in infrastructure automation and scripting with Terraform, Python, or Bash. - Experience managing CI/CD pipelines and deployment automation. - Strong troubleshooting and incident management skills.
Nice to have
- Experience supporting large-scale SaaS or cloud-native platforms. - Familiarity with workflow orchestration technologies such as Conductor, Temporal, or Camunda. - Experience with Kafka, messaging systems, or event-driven architectures. - Knowledge of security best practices and cloud infrastructure hardening. - Open-source contributions or a strong systems engineering background.
Benefits and work setup
- Base salary range of $125,000–$250,000 USD, with variation based on skills, experience, scope, location, and market data. - Full-time remote position open to candidates located in Australia or EMEA. - 15–20% travel expected. - Comprehensive health coverage including medical, dental, and vision. - Flexible PTO. - Support for personal development.