Remote job
Senior Software Engineer, Machine Learning Infrastructure & Automation
Job details
About this role
Role overview Help a machine learning organization ship generative AI models faster by building the automation, infrastructure, and developer tooling that supports model development, testing, and deployment. This senior engineering role focuses on eliminating manual work, accelerating development cycles, and creating reliable systems that allow ML engineers to ship models and optimizations with confidence.
Responsibilities - Design, build, and maintain CI/CD pipelines that automate testing, validation, and deployment for ML models and inference services - Dramatically reduce CI execution time through intelligent parallelization, caching, test selection, and efficient compute use - Develop automated model validation systems that detect quality regressions across models, GPU architectures, and configurations - Build continuous performance benchmarking that catches regressions in inference latency, throughput, GPU utilization, and cost - Create automated checks for pricing, billing, API schemas, and deployment readiness before production release - Extend AI-powered and agentic coding workflows that automatically diagnose CI failures, identify regressions, and propose fixes - Strengthen deployment reliability with automated safeguards, verification, rollback mechanisms, and monitoring - Identify and eliminate repetitive engineering toil by building tools and automation that free up ML engineers
Requirements - 5+ years of software engineering experience with strong Python skills and a background in production infrastructure and developer tooling - Hands-on experience designing and operating CI/CD systems using GitHub Actions or comparable technologies - Deep understanding of automated testing, build systems, dependency management, caching, and parallel execution - Experience with containerized workloads, Docker, and cloud infrastructure - Ability to design reliable distributed systems and debug complex infrastructure failures - Strong grasp of observability practices including logs, metrics, tracing, and automated alerting
Nice to have - Experience with ML infrastructure, PyTorch, GPU workloads, or model-serving systems - Familiarity with NVIDIA GPU architectures and multi-GPU environments - Background building automated inference benchmarks or ML quality evaluation frameworks - Experience with agentic coding tools or building custom AI engineering agents - Experience optimizing CI/CD pipelines at scale, including distributed test execution and ephemeral compute - Background developing internal developer platforms or infrastructure-as-code tooling
Benefits and work setup - Health, dental, and vision insurance for U.S. employees - Regular team events and offsites - Emphasis on learning, growth, and challenging technical work