Remote job
ML Ops Engineer
Job details
About this role
Role overview Take ownership of the machine learning platform at a fast-growing performance marketing company that connects consumers with new brands through thoughtful gifting experiences. Reporting to the Data Science Manager, this role partners closely with data scientists and product developers to productionize models, harden pipelines, and scale ML infrastructure as the business grows. The work spans training, inference, feature management, and observability across a high-volume, cash-flow-positive operation.
Responsibilities - Productionize batch and real-time training and inference, establish CI/CD for models, and define data and versioning practices plus model governance. - Centralize feature generation using feature store patterns, manage a model registry with metadata, and streamline deployment workflows. - Implement monitoring for data quality, drift, model performance, latency, and overall pipeline health with clear alerting and dashboards. - Refactor research code into reusable components and enforce repo structure, testing, logging, and reproducibility standards. - Collaborate with data scientists, analysts, and engineers to convert prototypes into production systems and provide mentorship. - Drive the technical roadmap for ML platform capabilities and establish architectural patterns that become team-wide standards.
Requirements - 5+ years of ML Ops experience, including ownership of ML infrastructure for large-scale production systems. - Strong software engineering skills in Python, with extensive experience building automated pipelines that bring ML models to production. - Production experience on AWS, Databricks, and container platforms such as Docker with Kubernetes (EKS, ECS, or equivalent). - Infrastructure as code experience using Terraform or CloudFormation for managed, reviewable environments. - Hands-on work with MLflow, SageMaker, or comparable ML tooling on production pipelines. - Experience with large-scale batch and stream processing using PySpark, Glue, Dask, or Kafka, plus familiarity with real-time endpoints, batch scoring, and feature stores. - Exposure to model governance, compliance, and secure ML operations, with strong communication across data and engineering teams.
Benefits and work setup - Competitive compensation with flexible remote work. - Unlimited responsible PTO. - Opportunity to join a growing, cash-flow-positive company with direct impact on revenue, scale, and growth.