Remote job
Tech Lead - MLOps & Infrastructure
Job details
About this role
Role overview
This technical leadership role focuses on designing and implementing production-grade MLOps infrastructure for life sciences AI workloads. The position owns end-to-end ML pipeline architecture, infrastructure-as-code automation, and full model lifecycle management while serving as the overall Tech Lead driving architectural decisions and technical standards. The work centers on building a reliable, scalable, and compliant ML platform that meets validation and reproducibility requirements for healthcare and life sciences use cases.
Responsibilities
- Architect and implement end-to-end MLOps pipelines spanning CI/CD for model training, evaluation, approval, and deployment with full audit trails - Design and deploy infrastructure-as-code using Terraform for all ML platform resources - Build automated training jobs on Amazon SageMaker, covering hyperparameter tuning, distributed training, and spot optimization for life sciences workloads - Establish CI/CD for ML artifacts including model versioning, container builds, integration testing, and staged rollouts gated by validation checks - Design model registry and artifact management supporting governance, reproducibility, and 21 CFR Part 11 compliance - Implement monitoring, drift detection, prediction quality tracking, automated retraining triggers, alerting, and auto-scaling for training and inference workloads - Define MLOps best practices, coding standards, and architectural patterns while leading architecture reviews, mentoring engineers, and coordinating with platform, security, and quality teams
Requirements
- Proven experience architecting production-grade MLOps pipelines, ideally in regulated environments such as healthcare or life sciences - Strong hands-on expertise with Amazon SageMaker, including hyperparameter tuning and distributed training - Deep proficiency with Terraform and infrastructure-as-code patterns for ML platforms - Demonstrated ability to build CI/CD systems for ML artifacts with versioning, containerization, and gated rollouts - Knowledge of model governance, reproducibility standards, and regulatory frameworks such as 21 CFR Part 11 - Experience implementing model monitoring, data drift detection, and automated retraining pipelines - Track record of technical leadership through architecture reviews, mentoring, and cross-team coordination on networking, security, and compliance
Nice to have
- Familiarity with HCLS AI validation and reproducibility standards - Background working alongside IT security and quality teams on regulated platform initiatives
Benefits and work setup
- People-first culture with a stated emphasis on individual growth and continuous learning - Servant leadership approach where management focuses on clearing roadblocks rather than micromanaging - Flat organizational structure with direct influence on the technical roadmap and client outcomes - Encouragement to learn from mistakes as part of an experimental, fast-moving engineering culture - Commitment to building diverse teams and inclusive hiring practices