Remote job
Research Scientist / Engineer – Performance Optimization
Job details
About this role
Role overview This position centers on making large multimodal models fast at every layer of the stack, from training through production deployment. The work combines low-level GPU optimization with architectural tuning for transformer-based systems, ensuring models run efficiently on modern accelerators without sacrificing output quality. It is a deeply technical role for someone who lives in profilers, kernels, and hardware-aware design.
Responsibilities - Profile and optimize GPU, CPU, and accelerator code paths to push utilization higher and latency lower. - Write high-performance PyTorch, Triton, and CUDA, including custom operations where off-the-shelf kernels fall short. - Design and ship fused kernels that take advantage of tensor cores and platform-specific hardware features. - Tune model architectures and implementations for distributed, multi-node production deployment. - Build the monitoring, analysis, and automation tooling that prevents performance regressions over time. - Research and apply state-of-the-art optimization techniques targeting transformer internals.
Requirements - Expert-level proficiency in Triton and CUDA, with a track record of serious GPU optimization work. - Strong PyTorch skills, including kernel development and custom-operation authoring. - Hands-on experience with profiling tools such as NVIDIA Nsight, the PyTorch profiler, and bespoke instrumentation. - Deep understanding of transformer architectures and attention mechanisms.
Nice to have - Familiarity with compilers and exporters such as torch.compile, TensorRT, ONNX, or XLA. - Experience tuning inference workloads for latency and throughput under real-world traffic. - Deep knowledge of the Triton compiler and kernel fusion techniques. - Comfort with warp-level intrinsics and advanced CUDA optimization patterns.
Benefits and work setup The role is positioned as an equal-opportunity position with the employer described as an equal opportunity employer; specific compensation, benefits, and remote/hybrid policies were not provided in the source material.