Remote job
Senior Infrastructure Engineer
Job details
About this role
Role overview
Own and evolve cloud infrastructure that supports large-scale customer-facing products and the engineering organization behind them. You will take responsibility for infrastructure products or subsystems from design through production operation, while improving reliability, security, scalability, performance, cost efficiency, observability, and developer productivity. The role requires both hands-on engineering and technical leadership across infrastructure, platform, backend, security, and product engineering teams.
Responsibilities
- Define goals, success measures, technical designs, and delivery plans for infrastructure work, managing risks and addressing important edge cases without unnecessary complexity. - Design, implement, operate, and improve secure, resilient, high-performing, and cost-conscious cloud infrastructure. - Build production software and developer tooling that reduces operational toil and improves the daily experience of engineering teams. - Manage infrastructure through code and configuration, primarily using Terraform or equivalent practices. - Partner with product engineers to design services for scale and resolve ambiguous technical requirements with stakeholders. - Participate in incident response, systematic debugging, monitoring improvements, performance tuning, and reliability initiatives. - Apply a security-focused approach to implementation and code review, actively identifying vulnerabilities and operational risks. - Mentor engineers through code reviews, pairing, design feedback, and knowledge sharing, while facilitating collaboration beyond engineering.
Requirements
- Six to ten years of experience in infrastructure, platform, or backend engineering in predominantly cloud-based environments; AWS experience is preferred. - Demonstrated end-to-end ownership of a significant product or subsystem, including design, delivery, production operation, and ongoing improvement. - Deep expertise in one or two infrastructure areas, with enough breadth to contribute across distributed systems and cloud platforms. - Strong hands-on experience with Terraform or comparable infrastructure-as-code tooling. - Solid knowledge of networking, load balancing, containerization, Kubernetes or EKS, and distributed-systems fundamentals. - Production programming experience in Go, Python, or a comparable language. - Hands-on Redis or ElastiCache operations, including sharding, failover, eviction policies, and scaling strategies. - Familiarity with Prometheus, Grafana, OpenTelemetry, or similar observability tools, as well as incident management, testing, source control, and code review.
Benefits and work setup
This is a fully remote position with team members working from many countries. Applicants must be authorized to work from their home location; visa sponsorship is not available, and a limited number of countries may be excluded for regulatory or security reasons.