Job details
About this role
Role overview A senior network engineering position focused on building and operating the high-performance fabric that underpins large-scale AI training and inference workloads. The engineer will design, deploy, automate, and maintain InfiniBand networking infrastructure supporting GPU-dense clusters, ensuring deterministic performance and predictable scaling as the platform grows.
Responsibilities - Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics powering AI/ML GPU clusters. - Administer NVIDIA Unified Fabric Manager (UFM) Enterprise for provisioning, telemetry, monitoring, and fabric health. - Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches, including firmware lifecycle management. - Troubleshoot fabric-level performance issues affecting NCCL, MPI, GPUDirect RDMA, and distributed training jobs. - Implement and validate fat-tree, Dragonfly+, Clos, and spine-leaf topologies for high-throughput AI environments. - Automate network provisioning through Python, Ansible, Git, REST APIs, and Infrastructure-as-Code practices. - Monitor infrastructure with UFM telemetry, Prometheus, and Grafana; lead incident response and capacity planning.
Requirements - 7+ years of experience operating complex network infrastructure at scale. - Deep expertise with InfiniBand networking, NVIDIA UFM, and large-scale GPU deployments. - Strong Linux networking fundamentals and hands-on troubleshooting of distributed workloads. - Proficiency with NCCL, RoCEv2, GPUDirect RDMA, and high-performance communication patterns. - Working knowledge of congestion control, adaptive routing, QoS, and traffic engineering. - Experience with automation tooling (Python, Ansible), version control (Git), and observability platforms.
Nice to have - Familiarity with Netris, Terraform, or comparable infrastructure tooling. - Exposure to multi-region backbone design and bare-metal provisioning systems. - Background operating HPC or GPU-dense environments, ideally in high-growth infrastructure settings.
Benefits and work setup - Anticipated base salary range of $170,000–$210,000 USD, plus discretionary bonus and equity. - Comprehensive medical, dental, and vision coverage for employees and dependents. - 401(k) matching (U.S.) and pension contributions (U.K.). - Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break. - Paid parental and family leave, plus four weeks of paid sabbatical after four years of service. - Annual learning and development allowance, wellness stipend, and work-from-home stipend. - Hybrid model with a minimum of two in-office days per week in San Francisco, Seattle, or NYC; fully remote considered outside hub locations. Occasional team and company offsites.