Remote job
Senior HPC Cluster Engineer
Job details
About this role
Role overview A senior engineering role on a GPU and high-performance networking team building core components of a large-scale cloud platform. The focus is on GPU computing, InfiniBand fabrics, and the KVM/QEMU virtualization stack, supporting multi-GPU, HPC environments. Day-to-day work spans low-level performance tuning, root-cause analysis, hardware integration, and automation for fault detection across complex infrastructure.
Responsibilities - Tune the performance of GPU clusters and InfiniBand networks for HPC and GPU-accelerated workloads. - Diagnose and troubleshoot root causes of issues affecting GPUs and InfiniBand fabrics, proposing corrective actions. - Integrate new GPU hardware into existing infrastructure through software stacks such as Kubernetes, QEMU, and KVM. - Build and enhance automation for proactive monitoring, detection, and resolution of issues in GPU and InfiniBand environments. - Configure and manage GPU devices and InfiniBand fabrics to ensure reliable, efficient operation. - Collaborate on hardware virtualization and device emulation work, balancing performance and security.
Requirements - 5+ years of professional experience in system-level software development with a focus on performance optimization and low-level programming. - 3+ years of hands-on Linux systems experience covering administration, troubleshooting, and performance tuning. - Deep understanding of server architecture, including PCIe devices, NICs, the Linux OS and kernel, and HPC systems. - Strong proficiency in one or more performance-oriented languages such as C, C++, Go, or Python.
Nice to have - Experience with end-to-end GPU testing in cluster environments using InfiniBand networking. - Track record of analyzing and optimizing HPC workloads such as simulations, data analysis, or AI/ML pipelines. - Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication. - Background in software-defined networking and HPC cluster networking. - Working knowledge of QEMU/KVM virtualization and managing virtualized environments. - Experience with deep learning frameworks like PyTorch and TensorFlow and their integration with HPC systems. - Familiarity with collective communication libraries such as MPI and NCCL for distributed computing.
Benefits and work setup - Competitive compensation package. - Opportunities for career growth and continued learning. - Flexibility and meaningful ownership over technical direction. - Collaborative culture oriented around innovation. - Work on infrastructure projects with significant impact in the AI space. - International team environment.