Remote job
Site Reliability Engineer
Job details
About this role
Role overview
This is a hands-on Site Reliability Engineer role focused on improving the reliability, performance, and operational maturity of a recently migrated AWS production environment. The position owns well-scoped reliability projects end to end and contributes to larger cross-functional initiatives, partnering closely with product engineers to build resilient systems. It is a remote-first U.S.-based role with periodic in-person team and company-wide gatherings.
Responsibilities
- Operate, maintain, and improve production infrastructure in AWS, including Terraform-managed environments and Kubernetes/EKS workloads. - Improve dashboards, alerts, and service-level indicators using New Relic or comparable observability platforms. - Investigate production issues, identify root causes, and implement durable fixes while participating in blameless postmortems and follow-up actions. - Participate in the shared 24/7 on-call rotation and contribute to effective incident response. - Improve application resilience using established patterns such as timeouts, retries, queuing, backpressure, and idempotency. - Maintain and improve CI/CD pipelines and deployment workflows using tools such as GitHub Actions and CircleCI. - Automate repetitive operational tasks, maintain runbooks and system documentation, and contribute to capacity planning, performance testing, and production-readiness reviews.
Requirements
- Approximately 5+ years of related experience in software engineering, infrastructure, systems engineering, SRE, platform engineering, DevOps, or equivalent practical experience. - Hands-on experience operating production workloads in AWS and building or maintaining infrastructure using Terraform or similar infrastructure-as-code tooling. - Experience with New Relic, Datadog, or another modern observability platform, and with troubleshooting production incidents in an on-call rotation. - Experience building or maintaining CI/CD pipelines, plus software development or scripting ability to read, debug, and modify application or automation code. - Working knowledge of Linux, networking, distributed systems, and relational databases. - Strong written and verbal communication skills, ability to articulate root causes and tradeoffs, and a habit of automating repetitive work while improving system reliability.
Nice to have
- Experience with Ruby or Ruby on Rails, PostgreSQL administration or performance tuning, Kubernetes and Amazon EKS, Redis, OpenSearch, or Amazon RDS. - Familiarity with SLOs, SLIs, error budgets, capacity modeling, or load testing, plus experience with payments, fintech, or other regulated systems and SOC 2 or similar security and compliance programs.
Benefits and work setup
- Remote-first role based in the U.S., with multiple company-wide and team-specific in-person gatherings throughout the year.