← Back to jobs

Remote job

Inference Infrastructure Architect (Remote)

AI Engineer Full-time Permanent China

Job details

Not specified Salary
China Eligibility
Senior Experience
Full-time Employment

About this role

Role overview An Inference Infrastructure Architect is needed to operate and scale a globally expanding bare-metal GPU fleet powering large-language-model serving at production scale. The role sits at the foundation layer, owning the path from hardware and Kubernetes through to an OpenAI-compatible inference endpoint, and shaping the architecture for both multi-tenant serverless traffic and dedicated enterprise deployments.

Responsibilities - Operate and expand the fleet, maximizing useful inference throughput per GPU-dollar while meeting latency and reliability SLOs and continuously reducing cost per token. - Build and run serverless serving pools using vLLM and SGLang with continuous batching, prefix caching, low-precision formats (FP8, FP4, INT4), and MoE expert parallelism for the head of the open-weight catalog. - Design the fleet layer on the Kubernetes Gateway API with KV-cache-aware routing, prefill/decode disaggregation, KV tiering, and multi-LoRA serving so hundreds of customer adapters share a base pool. - Stand up the bare-metal platform: Kubernetes with GPU Operator, topology-aware scheduling, LeaderWorkerSet, Kueue, Kata Containers for isolated tenants, and lifecycle management via OpenStack Ironic. - Manage weight logistics and elasticity with P2P model distribution (Dragonfly, safetensors streaming), warm pools, and autoscaling driven by inference metrics such as queue depth, KV occupancy, and TTFT. - Wire observability through DCGM and engine metrics into Prometheus and OpenTelemetry, and run capacity planning grounded in roofline math, batching curves, and utilization versus cost-per-token economics.

Requirements - Deep operational command of vLLM or SGLang: deployment, tuning, and upgrades, with fluency in TP/EP parallelism, quantization, KV-cache settings, prefix caching, and disaggregation. - Strong Kubernetes on bare metal skills, including GPU Operator, device plugins, node pools, topology-aware placement, gang scheduling, GitOps rollouts, and pod-placement debugging. - Performance engineering ability: reading engine and DCGM metrics, reasoning from roofline models, and sizing deployments quantitatively. - Proficiency in Python and Go for automation, plus real Linux, networking, and storage depth, with the judgment to recognize CUDA-bound problems and route them upstream. - Comfort working in both English and Chinese with open-source communities, and writing runbooks and design docs that others actually use.

Nice to have - Experience operating inference platforms at the scale of large hyperscalers or model labs, or as a heavy contributor to vLLM, SGLang, llm-d, NVIDIA Dynamo, Ray/KubeRay, Kata Containers, Volcano, Kueue, HAMi, Dragonfly, Mooncake, or Envoy. - Fine-tuning and RL infrastructure work, including LoRA pipelines, evaluation harnesses, and champion/challenger rollout. - Real-time voice latency experience with sub-second time-to-first-token budgets on live calls. - Multi-region deployments and data-residency compliance.

Benefits and work setup - Founding-seat influence on a greenfield inference platform, including a say in team hiring and operating norms. - Remote position based in mainland China with no relocation required, working asynchronously with a global team; visa sponsorship and additional hiring entities are available for future relocation. - Open-source-first culture with upstream contribution as part of the role and supported conference travel.

Skills detected in the listing

PythonGoKubernetesLLM
Detected Sep 16, 2026
Last verified Sep 16, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight