Remote job
Platform / Reliability Engineer
Job details
About this role
Role overview A platform and reliability engineering role focused on keeping a voice platform available at 99.99% while scaling into new regions and absorbing roughly tenfold growth in call volume. The work spans infrastructure, incident response, and the practices that keep a fast-growing system dependable.
Responsibilities - Operate and evolve the platform to consistently meet a 99.99% availability target - Scale infrastructure across additional regions to support large increases in call volume - Build and improve observability, alerting, and runbooks for production systems - Lead incident response and drive follow-up work to prevent repeat outages - Partner with application teams to make reliability a default rather than a special effort
Requirements - Substantial experience operating high-availability production platforms - Strong knowledge of cloud infrastructure, ideally with multi-region deployments - Familiarity with SRE practices such as SLOs, error budgets, and on-call rotations - Comfort scaling distributed systems through rapid, large-volume growth - Ability to balance reliability work with feature delivery in a fast-moving environment