Remote job
Machine Learning Platform Engineer
Job details
About this role
Role overview An AI-focused engineering team is hiring an ML Platform Engineer to build and operate the infrastructure behind its AI capabilities. The role spans model training, evaluation, deployment, inference, observability, and continuous improvement, partnering with researchers and product engineers to turn evolving model requirements into reliable, cost-efficient production systems.
Responsibilities - Build and operate the ML infrastructure and platforms powering AI products - Design systems for model training, evaluation, deployment, inference, and experimentation - Build and optimize model serving and inference infrastructure for high-throughput, low-latency workloads - Develop reliable pipelines for data preparation, training, evaluation, model release, and continuous improvement - Build evaluation and benchmarking infrastructure to measure model quality, performance, and regressions - Build production observability, monitoring, tracing, and alerting for AI/ML workloads - Identify and resolve bottlenecks across the ML stack and continuously improve system performance
Requirements - Strong software engineering fundamentals with experience building production systems - Experience building ML infrastructure, platforms, or production ML systems - Hands-on experience with model deployment, inference, evaluation, or data pipelines - Solid grasp of distributed systems and system reliability - Proficiency with Python and either PyTorch or JAX - Experience with LLM serving infrastructure such as vLLM, SGLang, or TensorRT-LLM - Ability to write clean, maintainable, production-quality code in ambiguous, fast-moving environments
Nice to have - Cloud infrastructure and GPU performance tooling experience - Familiarity with vector databases and retrieval infrastructure - Experience with ML/data pipelines and workflow orchestration - Bias toward ownership, experimentation, and continuous improvement