← Back to jobs

Remote job

Senior Network Engineer - InfiniBand / UFM

lightningai

Other Full-time Permanent US

Job details

$170,000—$210,000 Salary
US Eligibility
Senior Experience
Full-time Employment

About this role

Role overview A senior network engineering position focused on building and operating the high-performance fabric that underpins large-scale AI training and inference workloads. The engineer will design, deploy, automate, and maintain InfiniBand networking infrastructure supporting GPU-dense clusters, ensuring deterministic performance and predictable scaling as the platform grows.

Responsibilities - Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics powering AI/ML GPU clusters. - Administer NVIDIA Unified Fabric Manager (UFM) Enterprise for provisioning, telemetry, monitoring, and fabric health. - Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches, including firmware lifecycle management. - Troubleshoot fabric-level performance issues affecting NCCL, MPI, GPUDirect RDMA, and distributed training jobs. - Implement and validate fat-tree, Dragonfly+, Clos, and spine-leaf topologies for high-throughput AI environments. - Automate network provisioning through Python, Ansible, Git, REST APIs, and Infrastructure-as-Code practices. - Monitor infrastructure with UFM telemetry, Prometheus, and Grafana; lead incident response and capacity planning.

Requirements - 7+ years of experience operating complex network infrastructure at scale. - Deep expertise with InfiniBand networking, NVIDIA UFM, and large-scale GPU deployments. - Strong Linux networking fundamentals and hands-on troubleshooting of distributed workloads. - Proficiency with NCCL, RoCEv2, GPUDirect RDMA, and high-performance communication patterns. - Working knowledge of congestion control, adaptive routing, QoS, and traffic engineering. - Experience with automation tooling (Python, Ansible), version control (Git), and observability platforms.

Nice to have - Familiarity with Netris, Terraform, or comparable infrastructure tooling. - Exposure to multi-region backbone design and bare-metal provisioning systems. - Background operating HPC or GPU-dense environments, ideally in high-growth infrastructure settings.

Benefits and work setup - Anticipated base salary range of $170,000–$210,000 USD, plus discretionary bonus and equity. - Comprehensive medical, dental, and vision coverage for employees and dependents. - 401(k) matching (U.S.) and pension contributions (U.K.). - Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break. - Paid parental and family leave, plus four weeks of paid sabbatical after four years of service. - Annual learning and development allowance, wellness stipend, and work-from-home stipend. - Hybrid model with a minimum of two in-office days per week in San Francisco, Seattle, or NYC; fully remote considered outside hub locations. Occasional team and company offsites.

Skills detected in the listing

PythonTerraform
Detected Sep 2, 2026
Last verified Sep 2, 2026
Original source: job-boards.greenhouse.io · application link requires Hidden Jobs Access

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight