Remote job
Senior Manager, Cluster Engineering & Deployment
Job details
About this role
Role overview Senior Manager, Cluster Engineering & Deployment runs the end-to-end process that converts delivered hardware racks into accepted, production-ready GPU clusters. The role owns network bring-up, fabric cabling verification against port maps, GPU node integration with the fabric, cluster-level validation and burn-in (including RCCL/collective performance), and the final acceptance gate that controls revenue. It is a schedule-critical leadership position at an AI compute cloud provider, accountable for deployment velocity across multiple concurrent builds.
Responsibilities - Own and evolve the cluster deployment playbook, covering staged bring-up, automated config push, link and optics validation, cabling verification against L1 port maps, and fault triage during deployment windows. - Lead deployment engineering across concurrent cluster builds through team leads and on-site engineers, coordinating daily with data center integration field teams and cabling vendors. - Drive deployment velocity engineering, reducing bring-up time per cluster through tooling, pre-staging, and defect-source elimination, and set the targets the team is measured against. - Own defect feedback loops to Network Engineering on design issues, to Layer One teams on cabling quality, and to vendors on hardware and optics RMA patterns, holding partners accountable to resolution. - Define spares, test equipment, and deployment tooling requirements per site and standardize them across all sites. - Own cluster validation end to end, including bandwidth and latency baselines, collective (RCCL) performance tests, burn-in criteria, and go/no-go acceptance gates, raising the bar as the fleet scales.
Requirements - 10+ years across network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery, including managing engineers in a field or deployment setting. - Hands-on fabric bring-up experience at scale, with deployments involving hundreds of switches and thousands of links. - Strong operational rigor, with a track record of building and enforcing playbooks, gates, metrics, and blameless defect loops. - Team leadership with schedule accountability across multiple concurrent builds or sites.
Nice to have - GPU cluster validation experience, including NCCL/RCCL benchmarking. - Automation skills (Python, Ansible) applied to deployment. - Optics and link-layer debugging depth. - Experience with acceptance testing as a commercial, revenue-linked gate.
Benefits and work setup - Stock options. - 100% employer-paid Medical, Dental, and Vision insurance for employees. - Company contributions to a Health Savings Account. - 100% employer-paid Short Term and Long Term Disability insurance. - Life insurance plus voluntary supplemental options, with additional coverage such as Pet and Legal insurance available. - Supplementary health benefits including discounted virtual healthcare appointments and serious illness support. - Flexible Spending Account, 401(k), and an Employee Assistance Program. - Flexible PTO, paid holidays, parental leave, and in-office perks.