Remote job
Forward Deployed Engineer, Private LLM Inference
Job details
About this role
Role overview A forward-deployed engineering role focused on bringing private LLM inference into client environments where regulated data, contractual restrictions, or volume economics rule out hosted model APIs. The work involves choosing models and quantisation, standing up serving layers on client GPUs or cloud tenancy, and proving latency, throughput, and quality against real traffic. The role ends in clean handoff, with the client's own team fully owning the stack afterwards.
Responsibilities - Select and benchmark open-weight models against actual client tasks, documenting quality trade-offs across quantisation choices. - Deploy inference on client GPUs or cloud using vLLM, SGLang, TGI, or MLX, tuning KV cache sizing, batching, and prefill behaviour to the traffic shape. - Build a gateway layer with OpenAI-compatible endpoints, auth, per-tenant quotas, and routing between local and hosted models. - Port client prompts, tool schemas, and retrieval pipelines to the local model, closing quality gaps with structured evals. - Run structured output and tool calling reliably on smaller models, with explicit fallback paths for non-compliance. - Measure time-to-first-token, tokens per second, and cost per task under sustained load, publishing results to the client. - Write capacity plans, upgrade paths, and model swap procedures for the client platform team. - Set up monitoring for silent failure modes such as quality drift, cache thrash, and GPU memory pressure. - Pair with client engineers until the stack is fully owned internally.
Requirements - 5+ years of production engineering with meaningful Python depth and systems-level understanding of GPU memory and network throughput. - Hands-on experience deploying and operating at least one LLM serving stack (vLLM, SGLang, TGI, llama.cpp, or MLX) under real traffic. - Working knowledge of quantisation, context-length trade-offs, and what changes when moving from a frontier API to a local 30B-class model. - Comfortable working inside client infrastructure with Kubernetes, Terraform, and the cloud they already run. - Writes benchmark reports that state method, caveats, and unfavourable numbers honestly.
Nice to have - Experience with RAG pipelines, rerankers, and embedding models running locally. - Apple silicon inference with MLX and unified memory sizing for edge or workstation deployments. - Prior forward-deployed, MLOps, or platform engineering background.