Remote job
Senior/Staff Software Engineer, ML Ops
Job details
About this role
Role overview Join a healthcare technology team applying AI to pathology as a Senior or Staff Software Engineer focused on MLOps. The position centers on designing, building, and scaling the machine learning infrastructure that powers enterprise AI systems, bridging research and production for diagnostic and drug development workflows. The role blends hands-on engineering with technical leadership, including architectural decisions, mentoring, and cross-team collaboration.
Responsibilities - Architect and build infrastructure and automation, spanning both AWS and on-premises environments, to support ML application development and deployment - Lead system design and architectural discussions for the MLOps platform, balancing performance, security, and compliance requirements - Research, evaluate, and integrate new MLOps tools, frameworks, and best practices through focused technical initiatives - Partner with machine learning engineers, data scientists, product engineering, and infrastructure teams to move models from research into production - Optimize ML workflows so that models are deployed and monitored efficiently and reproducibly - Automate ML operations, including CI/CD for models, feature engineering pipelines, and deployment strategies using Kubernetes, Airflow, and related orchestration tools - Champion engineering excellence through coding standards, design reviews, and mentorship of more junior engineers
Requirements - BS or Master's degree in Computer Science, Computer Engineering, Software Engineering, or a closely related field - 5+ years of software engineering experience for the Senior level, or 8+ years for the Staff level, with a focus on production-grade frameworks or applications - Strong skills building complex, multi-language systems and scalable backend architecture - Hands-on experience with Kubernetes and cloud platforms, preferably AWS - Familiarity with observability and monitoring tools such as Prometheus, Grafana, or Datadog - Solid grasp of DevOps principles and infrastructure-as-code, including Helm and Terraform - Proficiency in Python plus exposure to additional programming languages - Experience owning internal development platforms and serving internal customers
Nice to have - Experience with ML frameworks such as PyTorch or Scikit-learn - Familiarity with data workflow orchestration frameworks like Airflow or Kubeflow - Deeper expertise in MLOps practices, including model lifecycle management, feature stores, model monitoring, and CI/CD for ML - Exposure to streaming data processing with Kafka, Flink, or Spark Streaming - Understanding of security and compliance best practices for ML systems - Experience using AI coding assistants such as Copilot or Cursor in day-to-day development - Strong communication, collaboration, and leadership skills across technical and non-technical stakeholders
Benefits and work setup - Hybrid position based in Boston, Massachusetts, with a remote option potentially available for exceptional candidates - Relocation benefits are not offered for this role - Expected salary range of $127,500 to $195,500 at the Senior level and $146,250 to $224,250 at the Staff level, based on primary location and determined by experience, qualifications, and other job-related factors permitted by law - Employer is an equal opportunity employer with a policy against unlawful discrimination based on protected veteran status, disability status, and other categories covered by applicable federal, state, or local law