Remote job
HPC Cluster Architect
Job details
About this role
Role overview
A senior HPC Cluster Architect role focused on designing and delivering large-scale GPU compute clusters. The position owns end-to-end architecture across compute, networking, storage, and physical design, translating customer requirements into production-ready, commercially optimised deployments. It is a hands-on technical authority role for someone with deep HPC experience who wants to see designs go live.
Responsibilities
- Own end-to-end cluster architecture for large-scale NVIDIA GPU deployments, from initial customer requirements through rack layouts, BOM, power and cooling design, and production handover - Design high-performance network fabrics spanning compute (InfiniBand, RDMA, NVLink/NVSwitch), storage, and WAN, including topology, oversubscription models, and scaling strategies - Engage directly with OEMs and vendors to validate hardware configurations, review quotes, and balance technical soundness with commercial optimisation - Provide technical oversight during deployment and bring-up, supporting hardware validation, performance testing, and acting as escalation point for complex integration issues - Act as a senior technical leader across Solutions Architecture, Cloud Engineering, and data centre partners, contributing to standardised reference designs and growing the HPC engineering function
Requirements
- Proven experience designing and delivering HPC or AI software stacks at scale, including workload profiling, scheduler configuration (SLURM, PBS, or equivalent), MPI/NCCL tuning, and distributed training frameworks such as PyTorch, JAX, or DeepSpeed - Deep understanding of GPU software environments, including CUDA, cuDNN, NCCL, driver stacks, and the tooling needed to run large-scale AI training and inference reliably in production - Hands-on experience optimising AI and HPC workloads across multi-GPU and multi-node setups, covering profiling, bottleneck identification, and performance tuning at both application and infrastructure layers - Working knowledge of containerisation and orchestration in HPC/AI contexts: Docker, Kubernetes, NVIDIA GPU Operator, and container-native workload management - Background in an OEM, hyperscaler, neo-cloud, or enterprise/research HPC environment, with exposure to the full design-to-deployment lifecycle for GPU-accelerated workloads - Ability to produce clear technical documentation and architecture diagrams for both engineering and executive audiences, with confidence engaging customers, vendors, and internal teams as a technical authority
Nice to have
- Experience with large-scale cluster performance benchmarking (NCCL tests, MLPerf, or equivalent) and familiarity with expected outcomes across GPU generations and topologies - Exposure to MLOps tooling and AI platform layers, including experiment tracking (MLflow, W&B), model serving frameworks (Triton, vLLM), and pipeline orchestration (Kubeflow, Airflow) - Familiarity with InfiniBand and high-performance networking as it relates to distributed training performance, sufficient to engage credibly on topology and tuning decisions
Benefits and work setup
- Competitive salary with an annual discretionary bonus scheme - Employee wellbeing benefits and 25 days of holiday plus public holidays - Flexible working arrangements, with remote or hybrid options depending on role and location - Real ownership and autonomy, with the trust to take initiative and experiment - Clear career progression and growth opportunities within a fast-growing organisation - Collaborative, international culture built on trust, transparency, and ownership