Remote job
Software Engineer, Inference
Job details
About this role
Role overview
A large-scale inference systems role focused on how cutting-edge models get served, integrated, and kept reliably online. The work spans scheduling, fleet management, deployment pipelines, and reliability across clusters and hardware providers, keeping expensive GPU fleets productive while meeting internal service-level objectives. It is firmly a systems seat rather than a modeling seat, and fits an engineer comfortable with model serving and Kubernetes at meaningful scale.
Responsibilities
- Ship new model architectures by integrating them into the inference engine. - Collaborate with research, engineering, and infrastructure teams to optimize model efficiency and deployments. - Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows. - Automate, test, and maintain inference services to maximize uptime and reliability. - Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines. - Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.
Requirements
- Strong Python and system-architecture skills. - Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or comparable frameworks. - Background with queues, scheduling, traffic control, and fleet management at scale. - Solid experience with Linux, Docker, and Kubernetes, including orchestration, deployment, and scheduling. - Familiarity with Redis and S3-compatible object storage.
Nice to have
- Modern networking stacks including RDMA technologies such as RoCE, InfiniBand, or NVLink. - High-performance large-scale ML systems running on 100+ GPUs. - CUDA programming experience, plus familiarity with FFmpeg or multimedia processing.