Remote job
Inference Infrastructure Architect (Remote)
Job details
About this role
Role overview An Inference Infrastructure Architect is needed to operate and scale a globally expanding bare-metal GPU fleet powering large-language-model serving at production scale. The role sits at the foundation layer, owning the path from hardware and Kubernetes through to an OpenAI-compatible inference endpoint, and shaping the architecture for both multi-tenant serverless traffic and dedicated enterprise deployments.
Responsibilities - Operate and expand the fleet, maximizing useful inference throughput per GPU-dollar while meeting latency and reliability SLOs and continuously reducing cost per token. - Build and run serverless serving pools using vLLM and SGLang with continuous batching, prefix caching, low-precision formats (FP8, FP4, INT4), and MoE expert parallelism for the head of the open-weight catalog. - Design the fleet layer on the Kubernetes Gateway API with KV-cache-aware routing, prefill/decode disaggregation, KV tiering, and multi-LoRA serving so hundreds of customer adapters share a base pool. - Stand up the bare-metal platform: Kubernetes with GPU Operator, topology-aware scheduling, LeaderWorkerSet, Kueue, Kata Containers for isolated tenants, and lifecycle management via OpenStack Ironic. - Manage weight logistics and elasticity with P2P model distribution (Dragonfly, safetensors streaming), warm pools, and autoscaling driven by inference metrics such as queue depth, KV occupancy, and TTFT. - Wire observability through DCGM and engine metrics into Prometheus and OpenTelemetry, and run capacity planning grounded in roofline math, batching curves, and utilization versus cost-per-token economics.
Requirements - Deep operational command of vLLM or SGLang: deployment, tuning, and upgrades, with fluency in TP/EP parallelism, quantization, KV-cache settings, prefix caching, and disaggregation. - Strong Kubernetes on bare metal skills, including GPU Operator, device plugins, node pools, topology-aware placement, gang scheduling, GitOps rollouts, and pod-placement debugging. - Performance engineering ability: reading engine and DCGM metrics, reasoning from roofline models, and sizing deployments quantitatively. - Proficiency in Python and Go for automation, plus real Linux, networking, and storage depth, with the judgment to recognize CUDA-bound problems and route them upstream. - Comfort working in both English and Chinese with open-source communities, and writing runbooks and design docs that others actually use.
Nice to have - Experience operating inference platforms at the scale of large hyperscalers or model labs, or as a heavy contributor to vLLM, SGLang, llm-d, NVIDIA Dynamo, Ray/KubeRay, Kata Containers, Volcano, Kueue, HAMi, Dragonfly, Mooncake, or Envoy. - Fine-tuning and RL infrastructure work, including LoRA pipelines, evaluation harnesses, and champion/challenger rollout. - Real-time voice latency experience with sub-second time-to-first-token budgets on live calls. - Multi-region deployments and data-residency compliance.
Benefits and work setup - Founding-seat influence on a greenfield inference platform, including a say in team hiring and operating norms. - Remote position based in mainland China with no relocation required, working asynchronously with a global team; visa sponsorship and additional hiring entities are available for future relocation. - Open-source-first culture with upstream contribution as part of the role and supported conference travel.