Remote job
Senior DevOps Engineer
Job details
About this role
Role overview A senior platform engineer is needed to take full ownership of the production infrastructure behind a high-volume email testing and sending SaaS that runs across multiple AWS regions. Beyond keeping the existing cloud estate reliable, secure, and cost-efficient, the role will lead a strategic, multi-quarter effort to migrate a portion of workloads onto rented bare-metal and colocation facilities in a deliberately hybrid model. The engineer will shape the technical case, design the target platform, and drive the project from planning through cutover.
Responsibilities - Operate and evolve a multi-account, multi-region cloud infrastructure with a strong focus on automation, zero-downtime deploys, and cost tagging. - Maintain and extend Terraform modules and workspaces, GitHub Actions pipelines, and ECS-based blue/green deployments using Docker images. - Steward the reliability stack (monitoring, alerting, capacity planning) and the security baseline covering IAM, secrets management, and edge rules. - Build the business and technical case for which workloads should migrate off managed cloud services and which should remain. - Design the bare-metal and colocation landing zone, including compute, networking, storage, observability, and secrets management. - Plan migration waves with dependency maps, cutover and rollback plans, dual-run periods, and safe traffic shifting; coordinate vendors, timelines, and engineering teams through completion.
Requirements - 5+ years in senior DevOps or platform engineering roles with direct production ownership. - Deep hands-on expertise with AWS services such as VPC, ECS, IAM, RDS or ElastiCache, and core networking. - Terraform at scale, including reusable modules and remote state or Terraform Cloud. - CI/CD experience with GitHub Actions and automated production deployments using strategies such as blue/green. - Strong Linux networking fundamentals: DNS, TLS, load balancing, firewalls, and VPN or hybrid connectivity. - Demonstrated experience planning and executing migrations to on-premises, colocation, or private cloud environments, including TCO modeling and rollback design. - Fluency in shell scripting and at least one higher-level language such as Python, plus fluent written and spoken English.
Nice to have - Bare-metal and colocation operations tooling such as NetBox for IPAM, Ansible, image-based provisioning with Packer, and HAProxy or Nginx. - Replacing managed services with self-hosted alternatives for PostgreSQL HA, Redis or Valkey, Kafka, and OpenSearch. - Email infrastructure experience including MTA, SMTP, IP reputation, DKIM/SPF/DMARC, and platforms such as Halon. - Cloudflare for DNS, WAF, and Access; observability stacks like Prometheus, Grafana, or Loki. - Background with multi-region SaaS, EU data residency, a FinOps mindset, and reading Ruby or Go code.