Remote job
Staff ML Engineer – AWS Trainium & SageMaker
Job details
About this role
Role overview
This is a deeply technical ML engineering role focused on training and operating models on Amazon SageMaker powered by AWS Trainium, custom silicon built for large-scale model training. The work spans the full stack, from hardware and compiler behavior through PyTorch training code up to production SageMaker pipelines. The role is forward-deployed, embedded with client teams, and targets real production training workloads rather than notebook experiments.
Responsibilities
- Train and operate models on SageMaker with Trainium as the underlying compute - Write and optimize PyTorch training code with a real understanding of how it compiles and executes on Trainium, including NeuronCore architecture, compiler behavior, and memory and throughput tradeoffs - Diagnose training run issues that originate at the hardware or compiler layer rather than the model layer - Translate a request for a Trainium training job into a working, cost-aware production pipeline end-to-end - Tune distributed training runs for throughput and cost on SageMaker's training infrastructure - Work directly with client and internal engineering teams to scope and deliver production training workloads
Requirements
- Strong hands-on PyTorch experience, ideally including distributed or multi-device training - Production experience with Amazon SageMaker for training and/or inference - Comfort working close to the hardware layer, with the ability to debug issues that are about the accelerator rather than the model - Solid Python fundamentals and comfort operating in a client-facing, production engineering environment
Nice to have
- Prior experience with AWS Trainium or Inferentia via the Neuron SDK - A track record of picking up new hardware targets quickly when accelerator-specific experience is absent