Remote job
Sr. Machine Learning Engineer
Job details
About this role
Role overview
This is a senior, remote position focused on building and operating production machine learning and large language model systems from end to end. The work spans data preparation, evaluation design, scalable inference, and ongoing monitoring. It suits an engineer who treats ML systems as software systems with reliability and cost discipline.
Responsibilities
- Design and implement production ML services that support both batch and real-time traffic, with clear interfaces and service level agreements. - Build evaluation harnesses that combine offline and online signals, establish baselines, and iterate quickly with tight feedback loops. - Own model training and fine-tuning workflows and/or retrieval-augmented generation pipelines, instrumented with strong observability. - Partner with product and engineering teams to translate requirements into shipped, measurable systems. - Continuously improve reliability, latency, and cost of inference for production workloads.
Requirements
- At least five years of experience building AI/ML or data-driven systems in production. - Strong Python engineering skills, including packaging, testing, and performance tuning, plus experience with mainstream ML tooling. - Practical experience with LLM integrations such as prompting, tool use, and retrieval-augmented generation, and/or traditional ML pipelines. - Hands-on experience with cloud infrastructure on AWS, GCP, or Azure, and containerization using Docker. - Comfort owning systems end to end, including monitoring and incident response.
Nice to have
- Experience with vector databases and retrieval evaluation. - Experience with distributed compute frameworks such as Ray or Spark, and with feature stores. - Experience with MLOps tooling, including CI/CD for models, model registries, and drift monitoring.