Remote job
Lead Site Reliability Engineer
Job details
About this role
Role overview Lead the reliability, resilience, and recoverability of mission-essential cloud platforms by directing infrastructure-level disaster recovery (DR) drills, platform rebuild validation, and automated deployment processes. This is a hands-on lead SRE role supporting high-visibility, high-stakes systems in a fast-paced, regulated environment where uptime and recoverability are continuously tested.
Responsibilities - Drive full-lifecycle platform portability and DR drill execution, including validation of rebuild procedures and DR playbooks against a 48-hour recovery target. - Verify end-to-end data completeness, integrity, and accuracy during drill exercises, documenting results and remediation recommendations. - Identify exit-readiness gaps across infrastructure, deployment automation, monitoring, and data recovery, then drive corrective actions with engineering teams. - Design and maintain automated Infrastructure-as-Code workflows using Terraform, AWS CloudFormation, and standardized CI/CD pipelines. - Manage and optimize Kubernetes clusters and containerized (Docker) workloads, including provisioning, scaling, and reliability improvements. - Build observability solutions with CloudWatch, Datadog, or comparable tools, and participate in on-call rotations, root-cause analysis, and incident response.
Requirements - Bachelor's degree with 8–10 years of SRE, DevOps, cloud, or infrastructure engineering experience, or 12 years of relevant experience with a high school diploma. - Expert-level hands-on knowledge of AWS services across compute, networking, storage, IAM, and serverless components. - Strong experience with Infrastructure-as-Code (Terraform, CloudFormation) and CI/CD pipeline design, including progressive delivery mechanisms. - Deep Kubernetes administration and container orchestration skills, plus proficiency in Python, Java, C#, or Go for automation. - Demonstrated ability to debug complex failure modes such as cascading failures, network partitions, backpressure, and eventual-consistency issues. - Must be a U.S. citizen or green card holder and able to obtain a Public Trust clearance.
Nice to have - AWS DevSecOps Engineer, Solutions Architect, SysOps, or Developer certification, and/or Kubernetes CKA or CKAD. - Familiarity with Zero Trust security models and cloud security best practices. - Experience with GitLab, Jenkins, or similar CI/CD platforms. - Background in highly regulated environments (healthcare, finance, federal, defense) and large-scale legacy-to-cloud modernization programs. - Prior involvement in large-scale DR drills, continuity of operations (COOP), or portability and executable readiness assessments.