Remote job
Senior AI Infrastructure Engineer
Job details
About this role
Role overview
A senior individual contributor role focused on the core infrastructure that routes inference requests, caches results, and runs reliably inside customer-owned clouds. The position sits at the intersection of distributed systems, AI serving, and platform engineering, and is part of a team building enterprise AI tooling deployable directly into customer environments. The role is remote and full-time.
Responsibilities
- Design and own the per-request routing layer that selects the most appropriate model across a broad set of providers. - Build the semantic cache, automatic failover, and provider-arbitrage mechanisms that hold cost and latency down. - Make the stack run cleanly inside customer clouds through Kubernetes, GPU scheduling, autoscaling, and safe upgrades. - Instrument systems around cost-per-resolved-task, not just raw latency, so optimisation targets the right metric. - Operate with a measurement-first mindset: instrument before optimising, and validate assumptions with data.
Requirements
- Five or more years building high-throughput backend or infrastructure systems in production. - Strong Python and distributed-systems fundamentals, including comfort with async and streaming patterns. - Hands-on experience with Kubernetes and at least one major cloud provider, spanning AWS, GCP, or Azure. - A bias for measuring before optimising, with the engineering discipline to back it up.
Nice to have
- Experience with LLM inference, GPU scheduling, or production model serving. - Prior work on multi-tenant, self-hosted software deployed inside customer-controlled environments.
Benefits and work setup
- Remote, full-time engineering role. - Applications are reviewed directly by an engineer rather than a recruiter, with replies expected within a few business days.