Remote job
Principal Engineer, Model Optimizations
Job details
About this role
Role overview
A principal-level role owning the model-optimization discipline across a heterogeneous GPU fleet powering a production inference cloud. The position spans every layer of inference performance, from quantization strategy and kernel authorship to parallelism topology and speculative decoding, across both NVIDIA and AMD stacks. It is a technically strategic role setting methodology, tooling, and vendor relationships rather than a pure implementation seat.
Responsibilities
- Set the technical direction for model optimization across the fleet, deciding which techniques are invested in internally versus consumed from upstream - Own quantization end to end, including FP8, FP4/MXFP4, INT8, and weight-only schemes such as AWQ and GPTQ, along with calibration methodology and accuracy-budget enforcement - Drive performance for modern architectures at the kernel level, covering MoE routing, fused expert kernels, MLA and GQA attention variants, and long-context attention patterns - Lead speculative decoding efforts using draft models, EAGLE/Medusa-class methods, and n-gram and lookahead approaches, including acceptance-rate tuning - Author and tune kernels in CUDA, Triton, and CUTLASS on one vendor stack and HIP, Composable Kernel, hipBLASLt, or AITER on the other, knowing when not to write custom code - Choose parallelism layouts (tensor, pipeline, expert, and attention data parallelism) per model, GPU family, and traffic shape, and encode those decisions into repeatable systems - Build benchmarking and regression infrastructure that validates TTFT, ITL, throughput, and accuracy before any change ships - Make AMD a genuinely first-class inference target by driving upstream contributions to vLLM, SGLang, and TensorRT-LLM, and by partnering with both vendors on pre-silicon enablement and roadmap feedback - Mentor senior and staff engineers and represent the organization in upstream communities, conferences, and customer-facing technical discussions
Requirements
- Twelve or more years in performance-critical systems with substantial recent production experience optimizing large-language-model inference - Deep knowledge of GPU architecture and the inference performance model, including memory-bandwidth and compute bounds, arithmetic intensity, kernel launch overhead, and the distinct needs of prefill versus decode - Hands-on kernel-level experience on at least one vendor stack (CUDA, CUTLASS, and Triton, or ROCm, HIP, and Composable Kernel) plus clear ability to work across both - Practical quantization expertise, including the judgment to identify when a technique that benchmarks well will fail a customer's accuracy bar - Familiarity with the internals of at least one major serving engine such as vLLM, SGLang, or TensorRT-LLM
Nice to have
- Publications or patents in efficient inference, quantization, or GPU kernel design - Same-week or day-zero enablement for newly released frontier open models
Benefits and work setup
- Hybrid work model - Base salary range of $249,600 to $312,000, with eligibility for bonuses tied to company and individual performance - Equity compensation for eligible employees, including initial grants and an Employee Stock Purchase Program option - Reimbursement for relevant conferences, training, and education, plus access to a large library of professional development courses - Flexible time off, employee assistance resources, and locally tailored benefits depending on region