Remote job
Principal Infrastructure Engineer
Job details
About this role
Role overview
Own the most difficult infrastructure and reliability challenges behind a high-volume digital commerce and payments platform. This principal-level role spans cloud architecture, Kubernetes, databases, networking, performance engineering, incident response, and AI-assisted SRE tooling, with accountability from initial design through production operation and measurable business impact.
Responsibilities
- Set the technical direction for core infrastructure by identifying system limits, prioritizing improvements, and delivering changes that support rising traffic, data volume, and workload complexity. - Connect customer and business workflows to infrastructure behavior, tracing performance, reliability, and cost issues across applications, data stores, and platform services. - Build capacity models, run load and stress tests, diagnose bottlenecks, and validate gains in throughput, latency, saturation, recovery time, and cost per workload. - Design resilient cloud foundations covering accounts, identity, networking, service quotas, fault isolation, and multi-zone or multi-region operation. - Improve the Kubernetes platform through lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and dependable deployments. - Scale and tune managed relational databases, addressing query and index performance, connection pressure, replication, failover, storage, and capacity constraints. - Participate in on-call operations and major-incident recovery, using logs, metrics, and traces to make mitigation decisions and convert findings into durable fixes. - Develop AI-assisted tools for investigation, capacity analysis, runbook automation, and toil reduction with appropriate validation, access controls, and auditability.
Requirements
- Principal-level experience solving complex infrastructure, distributed-systems, or site-reliability problems in production. - Deep AWS expertise, including cloud architecture, networking, IAM, resilience, quotas, and infrastructure-as-code. - Strong Kubernetes knowledge covering cluster internals, scheduling, isolation, scaling, upgrades, and deployment reliability. - Hands-on experience operating and optimizing MySQL and PostgreSQL workloads on a managed relational database service. - Ability to write production code and infrastructure automation, debug under pressure, and reason across application, platform, network, and database layers. - A history of measurable improvements in availability, latency, throughput, capacity, recovery, cost, or operational effort. - Clear communication, technical judgment, willingness to challenge decisions constructively, and accountability for outcomes after deployment.
Benefits and work setup
- Full-time remote position. - The working environment emphasizes open-source solutions, practical experimentation, rapid but calculated decisions, ownership, simplicity, and continuous improvement. - The commonly used stack includes Go and Python, AWS, Kubernetes, MySQL, PostgreSQL, Git, and GitLab-based CI/CD.