Remote job
Senior Software Engineer - Infrastructure
Job details
About this role
Role overview
Shape the cloud foundation that every engineering team deploys on, with a mandate to make releases predictable, scaling automatic, infrastructure self-serve, and cost transparent. This is a senior platform and infrastructure role focused on building the paved road for a large Python/AWS codebase serving high-volume production traffic, where AI-driven code velocity raises the bar for safety and reliability.
Responsibilities
- Redesign the deployment architecture to contain blast radius, isolating services so a failure in one area cannot cascade across the platform. - Own the container footprint on ECS and EKS, including networking between services, and evaluate broader Kubernetes adoption for emerging AI workloads. - Lead capacity engineering for a spiky workload, moving from reactive scaling to data-derived floors computed ahead of demand with early drift detection. - Build a self-serve Terraform platform with guardrails so product teams can own standard infrastructure changes while senior engineers focus on the non-standard. - Stand up per-team cost attribution across cloud, observability, and AI spend so tradeoffs are visible at the team level. - Steward the Python runtime and dependency health of the monolith, including garbage collection, event loop contention, and framework or package upgrades that are commonly deferred.
Requirements
- 5 or more years in platform, infrastructure, or SRE roles operating significant production traffic at scale. - Deep AWS expertise with hands-on production ownership of compute, networking, and identity (ECS, EKS, RDS, VPC, IAM). - Strong infrastructure-as-code background, particularly Terraform, including designing self-serve platforms consumed by other engineering teams. - Practical experience with autoscaling, capacity engineering, and container orchestration on ECS and/or EKS. - A track record of making deployments safe and self-serve across an organization rather than within a single team. - Comfort setting technical direction, driving alignment across teams, and raising the baseline through architecture reviews, runbooks, and paved-road tooling.
Nice to have
- FinOps and cloud cost optimization experience. - Deeper Kubernetes and EKS expertise. - Observability tooling at scale, ideally with Datadog. - Background in healthcare or other regulated environments.