Remote job
AWS Trainium / NKI Kernel Expert
Job details
About this role
Role overview This remote, part-time consulting role focuses on evaluating and refining AWS Trainium and Neuron Kernel Interface tasks for AI workloads. The work requires deep familiarity with Trainium architecture and an understanding of how it diverges from conventional GPU programming paradigms.
Responsibilities - Review NKI kernel correctness and Trainium-specific development patterns - Evaluate CUDA-to-NKI kernel migrations for idiomatic Trainium implementation - Assess performance optimization and benchmarking on Trainium hardware - Review memory management across the SBUF, PSUM, and HBM hierarchy - Analyze tile-based computation and DMA scheduling decisions - Evaluate cross-platform numerical correctness between CUDA/Triton and NKI - Identify Trainium-specific bottlenecks and surface optimization opportunities - Provide structured written technical feedback and quality assessments
Requirements - Two or more years of hands-on experience developing or optimizing kernels with the Neuron Kernel Interface - Experience working with AWS Trainium and/or Inferentia2 hardware - Strong grasp of tile-based computation, SBUF/PSUM/HBM memory hierarchy, partition dimension constraints, and DMA orchestration - Ability to evaluate CUDA-to-NKI migrations critically - Background profiling and optimizing workloads on Trainium - Understanding of numerical differences across GPU and Trainium backends
Nice to have - Experience with the AWS Neuron SDK or Neuron Compiler - CUDA or Triton kernel development background - Familiarity with NeuronCore-v2 architecture - Working with FP32, BF16, FP8, or INT8 workloads - Benchmarking on Trn1 or Trn2 instances - Knowledge of nki.language, @nki.jit, or XLA custom calls - Background in technical evaluation, RLHF, or rubric-based assessment
Benefits and work setup Remote, part-time, project-based consulting engagement focused on AWS Trainium and NKI kernel engineering and evaluation.