← Back to jobs

Remote job

Forward Deployed Engineer, Private LLM Inference

AI Engineer Freelance Remote

Job details

Not specified Salary
Remote Eligibility
Not specified Experience
Freelance Employment

About this role

Role overview A forward-deployed engineering role focused on bringing private LLM inference into client environments where regulated data, contractual restrictions, or volume economics rule out hosted model APIs. The work involves choosing models and quantisation, standing up serving layers on client GPUs or cloud tenancy, and proving latency, throughput, and quality against real traffic. The role ends in clean handoff, with the client's own team fully owning the stack afterwards.

Responsibilities - Select and benchmark open-weight models against actual client tasks, documenting quality trade-offs across quantisation choices. - Deploy inference on client GPUs or cloud using vLLM, SGLang, TGI, or MLX, tuning KV cache sizing, batching, and prefill behaviour to the traffic shape. - Build a gateway layer with OpenAI-compatible endpoints, auth, per-tenant quotas, and routing between local and hosted models. - Port client prompts, tool schemas, and retrieval pipelines to the local model, closing quality gaps with structured evals. - Run structured output and tool calling reliably on smaller models, with explicit fallback paths for non-compliance. - Measure time-to-first-token, tokens per second, and cost per task under sustained load, publishing results to the client. - Write capacity plans, upgrade paths, and model swap procedures for the client platform team. - Set up monitoring for silent failure modes such as quality drift, cache thrash, and GPU memory pressure. - Pair with client engineers until the stack is fully owned internally.

Requirements - 5+ years of production engineering with meaningful Python depth and systems-level understanding of GPU memory and network throughput. - Hands-on experience deploying and operating at least one LLM serving stack (vLLM, SGLang, TGI, llama.cpp, or MLX) under real traffic. - Working knowledge of quantisation, context-length trade-offs, and what changes when moving from a frontier API to a local 30B-class model. - Comfortable working inside client infrastructure with Kubernetes, Terraform, and the cloud they already run. - Writes benchmark reports that state method, caveats, and unfavourable numbers honestly.

Nice to have - Experience with RAG pipelines, rerankers, and embedding models running locally. - Apple silicon inference with MLX and unified memory sizing for edge or workstation deployments. - Prior forward-deployed, MLOps, or platform engineering background.

Skills detected in the listing

PythonKubernetesTerraformLLM
Detected Oct 5, 2026
Last verified Oct 6, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight