Remote job
Staff Software Engineer, AI/ML Infrastructure
Job details
About this role
Role overview Staff Software Engineer on a Machine Learning Infrastructure team that powers every AI-driven experience across a large home services marketplace. The engineer will be a senior technical leader spanning two surfaces: an AI/ML platform covering LLM gateways, model training and serving, feature and data workflows, orchestration, and observability, plus an evaluation platform for traces, scorers, regression gates, and self-serve tooling. Expect to design, write, and ship significant systems yourself while shaping multi-quarter roadmaps and build-versus-buy decisions.
Responsibilities - Build and evolve core AI platform capabilities that let product teams develop, run, and scale generative-AI applications. - Develop scalable tools and infrastructure for applied scientists, including model training and serving systems, feature and data workflows, CI/CD, orchestration, deployment, and evaluation tooling. - Work hands-on across the stack, from backend services and execution infrastructure to integrations with AI models and supporting tools. - Partner with senior engineers to evaluate emerging AI infrastructure frameworks and translate advances into platform capabilities. - Drive projects to completion with a strong focus on measurable business impact, and contribute to incident response for production AI systems.
Requirements - 10+ years of professional software engineering experience with a record of staff-level technical leadership across multiple teams in production. - 3+ years building AI/ML infrastructure in production, including model serving, feature platforms, training or orchestration systems, or ML observability. - Hands-on experience with LLM-powered systems in production, covering areas such as LLM gateways or proxies, prompt and model lifecycle management, RAG or agentic architectures, and LLM cost optimization. - Experience building or operating AI evaluation systems such as offline eval pipelines, LLM-as-judge calibration, regression testing for model or prompt changes, or A/B measurement of AI features. - Deep background designing, building, and operating reliable distributed systems at meaningful scale, including on-call responsibility and incident leadership. - Demonstrated use of AI coding tools in day-to-day work and the ability to validate and refine AI-generated output.
Nice to have - Comfort evaluating fast-moving vendor and open-source AI tooling landscapes and making build/buy/adopt recommendations.
Benefits and work setup - Expected total cash compensation range of $212,500–$275,000 for U.S. candidates outside the San Francisco Bay Area, based on calibrated level, skills, and experience. - Reasonable accommodations available for candidates with disabilities during the application process.