Remote job
Senior MLOps Engineer
Job details
About this role
Role overview
This senior MLOps role focuses on building the production backbone for AI systems on Google Cloud. The engineer will architect infrastructure, pipelines and tooling that take models from notebooks and proofs of concept into resilient, observable, auto-scaling services, working closely with AI researchers, data engineers and backend teams across a cybersecurity organisation.
Responsibilities
- Architect and manage scalable GCP-based ML infrastructure using Vertex AI, Google Kubernetes Engine, Cloud Storage, Cloud Run, and GPU or TPU compute - Own the end-to-end deployment lifecycle for machine learning models, building high-throughput, low-latency inference services with containerisation and specialised serving frameworks - Build automated, reproducible pipelines for model training, testing, evaluation and deployment using tools such as Airflow, Vertex AI Pipelines and GitHub Actions - Implement monitoring for system health and ML-specific metrics including feature drift, prediction accuracy and data distribution shifts, enabling automated retraining triggers - Provide AI and research engineers with scalable training environments, optimised runtime infrastructure and standardised deployment templates - Collaborate with data engineers to integrate pipelines with feature stores, dataset versioning and stream or batch processing workflows - Lead the technical transition of prototypes and notebooks into resilient, secure, auto-scaling microservices
Requirements
- At least five years of hands-on experience designing, deploying and maintaining production ML workloads in cloud environments - Deep practical experience with GCP, including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM and VPC configuration - Expertise with containerisation using Docker and Kubernetes or GKE, plus specialised serving tools such as Triton, vLLM or MLflow - Proven track record with workflow orchestrators like Airflow or Vertex AI Pipelines and modern CI/CD tools such as GitHub Actions or ArgoCD - Solid experience managing cloud resources with Terraform - Proficiency in Python and SQL for scripting, automation, API development and data manipulation - Hands-on experience with logging, telemetry and drift detection using Grafana, Prometheus, GCP Cloud Monitoring or specialised ML observability frameworks
Nice to have
- Experience running large-scale LLM or deep learning inference and training workloads - GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect certification - Familiarity with feature stores such as Feast or Vertex AI Feature Store
Benefits and work setup
- Opportunity to work across cloud, data and AI disciplines within a cybersecurity-focused product organisation - Exposure to production AI workloads at enterprise scale with autonomy to shape MLOps practices