Remote job
Senior Infrastructure Engineer, SRE
Job details
About this role
Role overview A consumer fintech is hiring a Senior Infrastructure Engineer with a focus on Site Reliability Engineering to lead the reliability and operational evolution of a large-scale platform. Hundreds of services run in production, processing billions of transactions, ingesting multiple terabytes of data, and emitting hundreds of millions of logs per day. The role joins the Cloud Infrastructure team and partners with product engineering and internal support teams to mature the reliability practice.
Responsibilities - Build and improve the reliability and resiliency of large-scale systems and services. - Establish SLIs, SLOs, and error budgets for the most critical services and user journeys, then review them with the owning teams. - Own and evolve the disaster recovery strategy, including recovery objectives, failover and restore paths, and regular exercise drills. - Partner with product engineering teams so they can own and operate their own services with metrics that reflect real user experience. - Evolve the observability platform and standards across metrics, tracing, and logs, including paved-path instrumentation, alert quality, and observability cost. - Contribute to day-to-day Cloud Infrastructure work alongside a shared on-call rotation of roughly one week every six weeks.
Requirements - At least five years of hands-on cloud or infrastructure engineering experience, with substantial time on reliability and production operations at scale. - A track record of defining SLIs and SLOs for real production services and explaining what changed as a result. - Hands-on experience with an enterprise observability platform in production, with Datadog strongly preferred. - Comfort writing code in Python, Go, TypeScript, or a similar language for internal tooling, production debugging, and automation. - Production-level Terraform and AWS skills, with the ability to find and fix problems when production breaks. - Personal experience being on-call for services you helped build, with clear opinions on what makes an alert worth waking someone for.
Nice to have - Leading a reliability or observability modernization project end to end. - Running game days, chaos experiments, or DR exercises and fixing the problems they uncover. - Cutting observability spend without sacrificing coverage.
Benefits and work setup - Health, dental, and vision plans, competitive pay, 401k matching, and unlimited PTO. - In-office perks include daily lunch, snacks, coffee, and commuter benefits. - Salary range disclosed in the listing is $150,000 to $185,000 per year plus bonus and benefits.