Remote job
Senior AI Platform Engineer
Job details
About this role
Role overview
A senior platform engineer is needed to lead the design and operation of the machine learning infrastructure that powers a digital health platform focused on metabolic care. This role sits at the intersection of backend distributed systems, MLOps, and applied AI, with responsibility for the systems that train, deploy, and monitor models in production. It is a hands-on technical leadership position working closely with data, infrastructure, and ML engineering teams.
Responsibilities
- Architect and build scalable AI/ML systems for production, including backend services, microservices, and pipelines that maintain model accuracy at scale.
- Lead cross-functional initiatives end-to-end, covering scoping, timelines, dependencies, and alignment across Data, Infrastructure, and ML Engineering.
- Partner with ML engineers to optimize workflows for model training, real-time inference, monitoring, and troubleshooting.
- Serve as a subject matter expert on ML infrastructure, advising internal teams and external partners on best practices.
- Define and enforce SLAs for system performance, including latency, throughput, and resource utilization, while driving operational reliability.
- Build tooling for model management, continuous monitoring, and lifecycle automation, and mentor teammates to raise the team's overall capability.
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field with 5+ years of relevant industry experience.
- Familiarity with architectural patterns for large, distributed, high-scale ML applications; experience with LLM and Generative AI implementations is a strong plus.
- Solid grounding in MLOps, data structures, and software design principles.
- Proficiency in Python, plus experience with at least one additional language such as Java or Go.
- Hands-on experience deploying scalable machine learning models using Docker, Kubernetes, and microservices architecture.
- Familiarity with SQL or NoSQL databases and big data processing frameworks such as Spark is a plus.
Nice to have
- Experience implementing applications that use LLMs and Generative AI in production.
- Background working with large-scale distributed ML platforms and observability tooling.
Benefits and work setup
- Compensation range of $180,000 to $200,000.
- Remote-first global team with competitive compensation aligned with leading technology companies.
- Equity participation, unlimited vacation with manager approval, and 16 weeks of 100% paid parental leave for delivering parents (8 weeks for non-delivering parents).
- 100% employer-sponsored healthcare, dental, and vision for employees, with 80% family coverage, plus HSA and FSA options, and a 401(k) retirement plan.