Remote job
Inference Acceleration Engineer
Job details
About this role
Role overview
This role sits on a small runtime team that writes the kernels, scheduler, and serving layer for a managed GPU inference platform that ships into customer production every week. The mission is to widen an existing 10–20× performance lead over hyperscaler inference on each model release. It is a systems-level individual-contributor position spanning kernel engineering, long-context research, and shared on-call reliability.
Responsibilities
- Profile live customer workloads with Nsight, rocprof, and an internal tracer, then convert the identified bottleneck into a reproducible benchmark and a merged kernel improvement. - Write fused, attention-aware kernels in CUDA and Triton, plus ROCm/HIP for AMD MI300X, targeting H100, H200, B200, and B300 accelerators. - Own one model family end-to-end, from tokenizer through KV cache through speculative decoding, across every supported accelerator. - Design and ship next-generation long-context primitives such as paged KV cache, ring attention, and sliding-window cache eviction. - Co-author at least one external write-up per quarter, whether a blog post, a paper, or a kernel released to the open-source ecosystem. - Carry a shared on-call rotation with cluster SRE and customer engineering (roughly one week out of six).
Requirements
- Five or more years of systems-level performance engineering, including at least two years of deep GPU kernel work. - Strong CUDA plus at least one of Triton, CUTLASS, or HIP. - Comfort reading PTX and SASS when profiling demands it. - A track record of measurable speed-ups shipped to production or in widely used open-source projects (a PR or perf graph expected). - Deep transformer internals knowledge: attention variants, KV cache shapes, FP8/INT4/NF4 quantization, and speculative decoding. - Empirical research instincts: designing experiments, holding variables fixed, and recording numbers before iterating.
Nice to have
- Published kernels in vLLM, SGLang, TensorRT-LLM, or FlashInfer. - Co-appointment with a partner university.
Benefits and work setup
- Full-time position at IC4–IC6 level. - Located in San Francisco or remote within the US and EU. - Stack spans CUDA, Triton, ROCm/HIP, PyTorch, C++17/20, NCCL/RCCL, plus Linux perf, Nsight, and rocprof. - Six-person runtime team reporting to the head of runtime, with weekly production release cadence.