Remote job
AI/ML Research Engineer
Job details
About this role
Role overview Train and improve foundational AI models, from pretraining and fine-tuning to evaluation and production deployment. The role spans research and systems engineering, with work across text, code, multimodal, and agent-oriented applications.
Responsibilities - Train and fine-tune foundation models, including workflows involving human or AI feedback. - Plan and run distributed training across GPU clusters. - Build evaluations for reasoning, tool use, and safety. - Work with inference and compiler teams to optimize and deploy models. - Produce reusable research outputs such as model weights, papers, and reproducible training recipes.
Requirements - Strong understanding of transformer architectures, attention, and current language-model training methods. - Practical experience with PyTorch, distributed training frameworks, and GPU performance profiling. - Evidence of delivering production machine-learning models, substantial open-source work, or published research. - Systems knowledge covering concurrency, networking, and memory efficiency.
Nice to have - Experience with synthetic data, agent reasoning methods, or tool-use tuning. - Familiarity with Triton, CUDA kernel development, or model quantization formats.
Benefits and work setup - Remote-first work across distributed locations. - Emphasis on measurable outcomes, open-source contributions, and individual ownership.