Remote job
Machine Learning Systems Engineer
Job details
About this role
Role overview
Engineer the pipelines and infrastructure that move machine learning models from experimentation into reliable production operation. The work includes data validation, model delivery, serving performance, and ongoing monitoring so teams can detect degradation and manage changes safely.
Responsibilities - Build and maintain training pipelines, feature workflows, and data validation stages. - Create deployment systems for model packaging, versioning, serving, testing, and rollback. - Monitor model performance and data drift, and implement triggers for retraining where appropriate. - Improve inference latency, throughput, and cost for batch and real-time workloads. - Coordinate with data engineering on feature stores and pipeline requirements. - Maintain experiment tracking, model registries, and lineage practices. - Document data dependencies and operational behavior for production handoff.
Requirements - Production experience building and operating ML systems, beyond research or experimentation. - Proficiency with frameworks such as TensorFlow, PyTorch, scikit-learn, or XGBoost. - Experience with MLOps platforms such as MLflow, Kubeflow, SageMaker, or Azure ML. - Understanding of model serving, deployment, monitoring, and the data dependencies of ML workflows. - Ability to optimize systems for operational reliability as well as model performance.
Benefits and work setup - Full-time role, based in Toronto or Montréal, or remote within Canada.