Remote job
GPU Kernel Engineer – CUDA, Triton & Accelerator Performance
Job details
About this role
Role overview A specialized consulting project invites experienced GPU kernel engineers to review, debug, and evaluate high-performance compute kernels used in modern AI workloads. The role centers on assessing whether kernel implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the underlying accelerator hardware. It suits engineers who enjoy working close to the silicon, tuning workloads, and pushing AI compute systems toward their performance limits.
Responsibilities - Review GPU and accelerator kernel implementations for correctness against reference outputs. - Evaluate numerical tolerance thresholds and floating-point precision behavior. - Analyze kernel benchmarks, performance profiles, and improvement claims for fairness and realism. - Identify performance bottlenecks across the memory hierarchy, occupancy, and compute throughput. - Audit kernel translations, operator fusion, and hardware migrations between frameworks. - Detect compilation, driver, memory, shape, and runtime issues and provide actionable technical feedback.
Requirements - 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels. - Strong working knowledge of at least two of CUDA, Triton, NKI/AWS Neuron, or Pallas/JAX. - Proficiency with profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers. - Deep understanding of memory bandwidth, GPU occupancy, shared memory, register pressure, memory coalescing, and bank conflicts. - Strong grasp of floating-point numerical correctness and tolerance thresholds. - Ability to distinguish software defects, environment problems, and genuine optimization challenges.
Nice to have - Experience across both NVIDIA GPU and custom accelerator ecosystems such as AWS Trainium or TPU. - Compiler engineering background or familiarity with MLIR, XLA, or intermediate representation lowering. - Contributions to GPU or ML kernel libraries such as cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls. - Experience with AI model evaluation, RLHF, or technical benchmark development.
Benefits and work setup Remote, part-time, project-based consulting engagement focused on GPU kernels, performance engineering, debugging, and technical evaluation.