Remote job
AI Research Engineer (Kernel & Inference Optimization) - 100% Remote Worldwide
Job details
About this role
Role overview This position drives innovation in how advanced AI models are served and run, especially on phones and other resource-constrained devices. The work spans the full stack of inference optimization, from writing low-level GPU kernels to designing serving pipelines that deliver high throughput, low latency, and minimal memory usage at scale.
Responsibilities - Design and deploy model serving architectures that balance throughput, latency, and memory footprint across cloud, mobile, and edge environments - Write and tune GPU compute kernels for mobile hardware, including custom shaders from scratch - Build end-to-end inference pipelines that take models from research prototypes into production on real devices - Engineer high-performance inference engines using techniques like tensor, pipeline, and expert parallelism to serve massive models on GPU clusters - Run controlled benchmarks in both simulated and live settings, tracking response latency, throughput, memory use, and error rates - Diagnose and fix bottlenecks such as suboptimal batching, network delays, and memory pressure on constrained hardware - Collaborate with cross-functional teams to integrate optimized serving frameworks into production stacks for edge and on-device applications
Requirements - Proven track record of low-level kernel optimization and inference work on mobile devices, with measurable improvements in latency, throughput, or memory footprint - Strong expertise writing GPU kernels for smartphone hardware and a deep understanding of model serving frameworks - Hands-on experience deploying end-to-end inference pipelines, from model optimization through on-device integration - Solid grasp of modern serving architectures and inference optimization techniques for low-latency, high-throughput, memory-efficient deployment - Familiarity with pruning, quantization, flash attention, KV cache, speculative decoding, and related acceleration methods - Understanding of the math and structure behind diffusion models and vision transformers
Nice to have - PhD in NLP, Machine Learning, or a closely related field, complemented by publications at top-tier conferences
Benefits and work setup - 100% remote role open to candidates worldwide - Globally distributed research team with high autonomy - Research-driven culture that values novel approaches and rigorous experimentation