Remote job
Director of Engineering, Cloud & Reliability
Job details
About this role
Role overview This director-level engineering role leads the cloud infrastructure and reliability function for a client experience platform serving appointment-based self-care businesses. The position owns the technical strategy for cloud, DevOps, and platform capabilities, and is responsible for transforming the infrastructure into a scalable, secure, highly available foundation. It is a fully remote position within the U.S.
Responsibilities - Define and execute the long-term infrastructure and security strategy so the platform remains scalable, reliable, and secure as usage grows. - Lead, mentor, and grow a team of infrastructure and reliability engineers, building a culture of operational excellence. - Lead migration to containerized, service-ready architecture on Amazon EKS. - Design and implement multi-region, hot-hot disaster recovery with automated failover, load balancing, and regular testing. - Modernize CI/CD pipelines and scale Terraform and GitOps practices to enable self-service infrastructure. - Strengthen monitoring, alerting, and incident response using observability tooling to detect and resolve issues early. - Drive architectural decisions across cloud, security, and operations, including scaling the database architecture beyond current RDS limits. - Partner with Product, Services, and Data Engineering teams and maintain compliance with PCI, HIPAA, and SOC 2.
Requirements - 10+ years of experience in cloud infrastructure and platform engineering, including 5+ years leading infrastructure teams of 5-15 engineers. - Deep experience with AWS services such as EKS, RDS, VPC, and IAM, plus infrastructure automation using Terraform. - Proven track record migrating applications to Kubernetes/EKS and implementing container-based architectures. - Strong background in security practices and compliance frameworks including PCI, HIPAA, and SOC 2, along with security automation. - Experience transforming CI/CD pipelines, implementing GitOps, and enabling rapid software delivery. - Expertise with monitoring and observability tools such as Datadog, Prometheus, and Grafana, with the ability to design thoughtful alerting. - Demonstrated ability to partner with Product Engineering, Data Engineering, and Platform Services teams to enable their initiatives.
Benefits and work setup - Fully remote within the United States. - Compensation varies by U.S. metro area; example ranges are provided for NYC, SF Bay Area, and Seattle, with a separate Canadian range of CAD $227,000-$283,800. - Benefits include performance bonus eligibility, equity participation, comprehensive medical coverage with employer-paid premiums, dental, vision, life, and disability insurance, an HSA contribution, 12 weeks of fully paid parental leave, flexible PTO plus company holidays, a home office setup allowance and monthly WFH reimbursement, and quarterly well-being stipends.