Remote job
Staff Machine Learning Infrastructure Engineer, Embedding Platform
Job details
About this role
Role overview This is a staff-level engineering role responsible for the infrastructure that produces, stores, and serves machine-learning embeddings at production scale. The position sits at the intersection of large-scale systems engineering and applied ML, supporting downstream teams that rely on embeddings for personalization, retrieval, and ranking. The work spans platform design, operational reliability, and technical leadership across partner teams.
Responsibilities - Architect and evolve the embedding platform, including training pipelines, feature stores, and low-latency serving systems. - Design and operate vector retrieval and similarity-search infrastructure that meets scale and latency targets. - Partner with applied scientists and ML engineers to onboard new embedding use cases and unblock production deployments. - Set technical direction, perform design reviews, and mentor engineers across the organization on ML systems best practices. - Own reliability, observability, and performance of embedding services, including on-call responsibilities. - Evaluate and integrate new technologies such as vector databases, hardware accelerators, and orchestration frameworks where they provide clear value.
Requirements - Extensive experience building and operating large-scale distributed systems, ideally with ML serving or data infrastructure focus. - Strong proficiency in languages such as Python, Scala, Java, or Go and comfort with modern ML and data tooling. - Deep understanding of model training pipelines, feature engineering, and the trade-offs between batch and online inference. - Familiarity with embedding techniques, approximate nearest neighbor search, and retrieval-augmented architectures. - Demonstrated technical leadership and ability to drive cross-team initiatives without direct authority.