Remote job
ML Infrastructure Engineer
Job details
About this role
Role overview An infrastructure engineering role focused on model serving at scale, where small percentages of throughput translate directly into whether the business is profitable. The work covers runtime selection, quantisation, evaluation and reproducible benchmarking, and the layer is treated as a first-class engineering surface rather than a deployment detail. Results land directly in what customers pay and the latency they observe.
Responsibilities - Own the model serving stack end to end, including runtime selection, configuration and upgrade path for each workload class - Tune throughput and latency through continuous batching, KV cache strategy, paged attention, tensor and pipeline parallelism, and speculative decoding where it pays off - Run quantisation work such as FP8, AWQ and GPTQ, and build the evaluation harness that quantifies quality trade-offs - Maintain reproducible benchmarking across model families and hardware generations, including the methodology write-ups behind published results - Profile and eliminate bottlenecks across the request path, covering GPU kernels, host-side overhead, batching and network - Partner with the GPU systems team on node-level tuning and with platform engineering on autoscaling and scheduling - Keep a clear eye on serving-side cost per token and make latency versus cost trade-offs explicit
Requirements - Shipped and operated at least one production model serving stack at meaningful scale - Strong programming ability in Python and enough comfort in CUDA or C++ to read a kernel and reason about it - Real experience with benchmarking and evaluation methodology, including knowing how to avoid fooling yourself with a flattering measurement - Familiarity with quantisation and its quality implications, rather than treating it as a switch to flip - Ability to communicate performance results to non-specialists without hand-waving or drowning them
Nice to have - Open-source contributions to a serving or inference project - Experience with fine-tuning workloads as well as inference - Exposure to multi-modal or embedding workloads alongside text generation - Familiarity with public inference benchmarking suites and their failure modes
Benefits and work setup - Remote-eligible, full-time engineering role with competitive compensation