← Back to jobs

Remote job

Inference Acceleration Engineer

AI Engineer US, EU

Job details

Not specified Salary
US, EU Eligibility
Senior Experience
Not specified Employment

About this role

Role overview

This role sits on a small runtime team that writes the kernels, scheduler, and serving layer for a managed GPU inference platform that ships into customer production every week. The mission is to widen an existing 10–20× performance lead over hyperscaler inference on each model release. It is a systems-level individual-contributor position spanning kernel engineering, long-context research, and shared on-call reliability.

Responsibilities

- Profile live customer workloads with Nsight, rocprof, and an internal tracer, then convert the identified bottleneck into a reproducible benchmark and a merged kernel improvement. - Write fused, attention-aware kernels in CUDA and Triton, plus ROCm/HIP for AMD MI300X, targeting H100, H200, B200, and B300 accelerators. - Own one model family end-to-end, from tokenizer through KV cache through speculative decoding, across every supported accelerator. - Design and ship next-generation long-context primitives such as paged KV cache, ring attention, and sliding-window cache eviction. - Co-author at least one external write-up per quarter, whether a blog post, a paper, or a kernel released to the open-source ecosystem. - Carry a shared on-call rotation with cluster SRE and customer engineering (roughly one week out of six).

Requirements

- Five or more years of systems-level performance engineering, including at least two years of deep GPU kernel work. - Strong CUDA plus at least one of Triton, CUTLASS, or HIP. - Comfort reading PTX and SASS when profiling demands it. - A track record of measurable speed-ups shipped to production or in widely used open-source projects (a PR or perf graph expected). - Deep transformer internals knowledge: attention variants, KV cache shapes, FP8/INT4/NF4 quantization, and speculative decoding. - Empirical research instincts: designing experiments, holding variables fixed, and recording numbers before iterating.

Nice to have

- Published kernels in vLLM, SGLang, TensorRT-LLM, or FlashInfer. - Co-appointment with a partner university.

Benefits and work setup

- Full-time position at IC4–IC6 level. - Located in San Francisco or remote within the US and EU. - Stack spans CUDA, Triton, ROCm/HIP, PyTorch, C++17/20, NCCL/RCCL, plus Linux perf, Nsight, and rocprof. - Six-person runtime team reporting to the head of runtime, with weekly production release cadence.

Skills detected in the listing

GDPRLLM
Detected Sep 28, 2026
Last verified Sep 28, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight