Remote job
Senior MLOps Engineer
Job details
About this role
Role overview Own the reliability, scalability, and automation of production machine-learning pipelines. This hands-on role spans cloud infrastructure and the ML lifecycle, evolving systems used by customers while delivering improvements incrementally and safely.
Responsibilities - Package Python pipeline services with Docker and deploy them on cloud infrastructure. - Design scalable, event-driven execution for data preparation, training, and prediction workloads. - Build and maintain automated build, test, and deployment workflows. - Improve observability through logs, metrics, and alerts. - Strengthen reproducibility, experiment tracking, model versioning, and deployment practices. - Modernize pipeline architecture while protecting live services and managing infrastructure costs.
Requirements - At least five years running production Python systems, with sound practices for testing, version control, code review, and CI/CD. - Practical Docker skills, including building and optimizing images and diagnosing containers in production. - Cloud experience across compute, storage, monitoring, and access control; the described environment uses AWS services. - Experience building automated CI/CD pipelines with GitHub Actions or an equivalent tool. - Working knowledge of both document and relational databases, including secure connections from containerized services. - Production ML lifecycle experience, including reproducible training, experiment tracking or model management, and deployment.
Nice to have - Familiarity with Python environment and configuration tools such as uv, Pydantic, Hydra, or OmegaConf. - Workflow orchestration experience, for example with Prefect or managed ML pipelines. - Familiarity with common ML libraries and geospatial tools or data. - Experience partnering with data scientists and platform or operations engineers.