Remote job
Senior Solution Engineer – GPU & AI Infrastructure
Job details
About this role
Role overview
This senior pre-sales architect role leads the design of large-scale GPU compute environments for enterprise AI and high-performance computing customers. The position sits at the intersection of business requirements and hardware execution, producing detailed design documentation for clusters built on the latest GPU architectures and high-speed fabrics. It is a hands-on technical leadership role combining systems engineering, customer engagement, and cross-team collaboration.
Responsibilities
- Author High-Level Design (HLD), Low-Level Design (LLD), and detailed Bill of Materials (BOM) documents for enterprise-scale GPU supercomputing clusters. - Lead technical engagements, translating complex AI training and inference workload requirements into production-ready infrastructure blueprints. - Architect solutions spanning bare-metal and Kubernetes-based orchestration, leveraging current-generation NVIDIA Blackwell GPUs and ultra-low-latency fabrics. - Run proof-of-concept validation, including fabric tuning, congestion mitigation, and thermal and power optimization during rollout. - Act as the bridge between enterprise clients, hardware vendors, and internal platform engineering teams, feeding market insights back into product planning.
Requirements
- 5+ years in solution architecture, systems engineering, or technical pre-sales focused on AI, HPC, or high-performance cloud infrastructure. - Bachelor's degree in Computer Science, Electrical Engineering, Systems Engineering, or equivalent practical experience. - Deep hands-on knowledge of NVIDIA HGX and DGX platforms, NVLink/NVSwitch fabrics, and Blackwell-class GPUs (B300, GB300NVL, GB200 NVL72/NVL36). - Expert-level familiarity with InfiniBand (Quantum-2/Quantum-X800, adaptive routing) and RoCE/RoCEv2 with Spectrum-X/Spectrum-4 Ethernet, including GPUDirect RDMA and Storage. - Proficiency deploying GPU workloads on Kubernetes (CNI, GPU Operator, RDMA plugin, MPI Operator) and bare-metal stacks (Slurm, Ansible, Terraform, PyTorch/NCCL tuning). - Understanding of high-density datacenter environments, liquid cooling approaches (direct-to-chip, CDU/liquid loop), and power delivery constraints for 100kW+ per rack; ability to produce rack diagrams and itemized BOMs. - Must be UK-based, with strong presentation and stakeholder communication skills.
Nice to have
- NVIDIA Certified Professional credentials such as AI Infrastructure (NCP-AII), AI Networking (NCP-AIN), or InfiniBand (NCP-IB). - NVIDIA Certified Associate or Professional in AI Workload Deployment and Cloud Native.
Benefits and work setup
- Remote-first working environment with flexibility and autonomy. - Four-day workweek (unless attending an event). - Uncapped holiday allowance and competitive compensation and benefits package. - Collaborative, inclusive culture described as valuing diversity, creativity, and innovation.