Remote job
Sr Machine Learning Engineer - AI
Job details
About this role
Role overview A senior engineering role focused on building, optimizing, and deploying small language models for edge and on-device experiences. The work spans the full ML lifecycle, from fine-tuning and quantization to MLOps pipelines and production monitoring under strict latency targets. It sits inside an applied AI team that supports a secure multi-cloud content governance platform.
Responsibilities - Fine-tune and train small language models using Hugging Face, TRL, and adapter methods such as LoRA, QLoRA, and PEFT. - Optimize models for inference through quantization, pruning, and knowledge distillation to reduce footprint and latency. - Deploy models to edge devices, mobile platforms, and local servers against defined latency and performance targets. - Build end-to-end MLOps pipelines covering data ingestion, training, deployment, and ongoing operations. - Monitor model accuracy, latency, and CPU/GPU utilization in production environments. - Evaluate model quality using standard benchmarking frameworks and custom evaluation suites.
Requirements - Hands-on experience training and fine-tuning small language models with Hugging Face and adapter-based techniques. - Practical knowledge of model optimization techniques including quantization, pruning, and knowledge distillation. - Experience deploying models to edge devices, mobile environments, and local servers. - Ability to build complete MLOps pipelines from data ingestion through deployment. - Skill in tracking model accuracy, latency, and hardware utilization in production.
Nice to have - Direct deployment experience on edge or mobile environments. - Familiarity with ONNX export and cross-platform inference. - Exposure to MLOps tooling such as experiment tracking, model registries, and CI/CD for machine learning.