← Back to jobs

Remote job

Senior HPC Cluster Engineer

Other Full-time Permanent Europe and UK

Job details

Not specified Salary
Europe and UK Eligibility
Senior Experience
Full-time Employment

About this role

Role overview A senior engineering role on a GPU and high-performance networking team building core components of a large-scale cloud platform. The focus is on GPU computing, InfiniBand fabrics, and the KVM/QEMU virtualization stack, supporting multi-GPU, HPC environments. Day-to-day work spans low-level performance tuning, root-cause analysis, hardware integration, and automation for fault detection across complex infrastructure.

Responsibilities - Tune the performance of GPU clusters and InfiniBand networks for HPC and GPU-accelerated workloads. - Diagnose and troubleshoot root causes of issues affecting GPUs and InfiniBand fabrics, proposing corrective actions. - Integrate new GPU hardware into existing infrastructure through software stacks such as Kubernetes, QEMU, and KVM. - Build and enhance automation for proactive monitoring, detection, and resolution of issues in GPU and InfiniBand environments. - Configure and manage GPU devices and InfiniBand fabrics to ensure reliable, efficient operation. - Collaborate on hardware virtualization and device emulation work, balancing performance and security.

Requirements - 5+ years of professional experience in system-level software development with a focus on performance optimization and low-level programming. - 3+ years of hands-on Linux systems experience covering administration, troubleshooting, and performance tuning. - Deep understanding of server architecture, including PCIe devices, NICs, the Linux OS and kernel, and HPC systems. - Strong proficiency in one or more performance-oriented languages such as C, C++, Go, or Python.

Nice to have - Experience with end-to-end GPU testing in cluster environments using InfiniBand networking. - Track record of analyzing and optimizing HPC workloads such as simulations, data analysis, or AI/ML pipelines. - Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication. - Background in software-defined networking and HPC cluster networking. - Working knowledge of QEMU/KVM virtualization and managing virtualized environments. - Experience with deep learning frameworks like PyTorch and TensorFlow and their integration with HPC systems. - Familiarity with collective communication libraries such as MPI and NCCL for distributed computing.

Benefits and work setup - Competitive compensation package. - Opportunities for career growth and continued learning. - Flexibility and meaningful ownership over technical direction. - Collaborative culture oriented around innovation. - Work on infrastructure projects with significant impact in the AI space. - International team environment.

Skills detected in the listing

PythonGoC++Kubernetes
Detected Sep 25, 2026
Last verified Sep 25, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight