Remote job
Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Job details
About this role
Role overview A Senior Applied Scientist role focused on efficient LLM and VLM inference and model optimization within an AI cloud platform. The position blends rigorous research with production engineering impact: turning frontier inference bottlenecks into well-scoped research problems, publishing credible work, and shipping the results into deployed inference capabilities. It is explicitly not a papers-only role, and close collaboration with machine learning engineers is expected throughout.
Responsibilities - Own focused research projects from hypothesis through experiment, ablation, prototype, and production handoff. - Prepare internal reports, technical blogs, or external papers when work is at publication quality. - Partner directly with machine learning engineers to turn research prototypes into usable production components. - Define and execute research programs in efficient LLM and VLM inference with measurable production impact. - Invent, evaluate, and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse and compression, long-context inference, MoE routing, and model/runtime co-optimization. - Build prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks, and collaborate on productionization. - Design rigorous evaluations covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token. - Mentor engineers and scientists on experimental design and model/system tradeoffs.
Requirements - PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field. - Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, or serving systems. - Strong hands-on coding ability in Python and PyTorch, with the ability to move from idea to experiment to prototype quickly. - Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs. - Strong experimental design skills, including ablations, baselines, metrics, statistical reasoning, and failure analysis. - Excellent written and verbal communication.
Nice to have - First-author publications at NeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, HPCA, or comparable venues. - Experience deploying ML models or inference optimizations in production. - Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals. - Experience with post-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation tied to inference quality or efficiency. - Open-source research artifacts, widely used benchmarks, technical blogs, or invited talks in efficient AI systems.
Benefits and work setup - Competitive compensation. - Career growth and learning opportunities in a fast-moving AI environment. - Flexibility and meaningful ownership over research direction. - Collaborative and innovative culture across international teams. - Opportunity to work on high-impact AI infrastructure projects.