Remote job
Member of Technical Staff, Inference
Job details
About this role
Role overview This position focuses on advancing a leading open-source inference engine used to serve large language and diffusion models at scale. The engineer will work at the core of the runtime, optimizing how models execute across diverse hardware and adapting the stack to emerging architectures such as mixture-of-experts, multimodal, and agentic systems. The work directly shapes the cost, latency, and reach of modern AI deployment.
Responsibilities - Design and implement runtime optimizations for transformer-based and diffusion model serving - Adapt the inference engine to new architectures including mixture-of-experts and multimodal workloads - Profile, debug, and resolve performance bottlenecks across heterogeneous hardware targets - Translate techniques from research papers into production-ready, maintainable code - Contribute to a complex, performance-critical ML systems codebase used by a broad community - Collaborate with open-source contributors and downstream integrators on systems improvements
Requirements - Bachelor's degree or equivalent experience in computer science, engineering, or a related field - Deep understanding of transformer architectures and their variants - Strong Python programming skills with hands-on experience working inside PyTorch - Practical experience with LLM inference systems such as vLLM, TensorRT-LLM, SGLang, or TGI - Ability to read research papers and implement novel model architectures and inference techniques - Demonstrated ability to ship performant, maintainable code and debug complex ML systems
Nice to have - Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving - Familiarity with reinforcement learning frameworks and algorithms applied to LLMs - Experience with multimodal inference across audio, image, video, and text - Prior contributions to open-source ML or systems infrastructure projects - Implemented core features in a major open-source inference engine - Built integrations with training frameworks such as verl, OpenRLHF, Unsloth, or LLaMA-Factory - Authored widely-shared technical writing or side projects on LLM inference
Benefits and work setup - Fully remote with hiring open worldwide - Timezone-flexible schedule with expected overlap with Pacific Time for critical syncs - Competitive salary and equity benchmarked to local market conditions - Visa sponsorship considered on a case-by-case basis - Location-appropriate benefits, including health coverage where applicable