← Back to jobs

Remote job

Research Scientist / Engineer – Training Infrastructure

AI Engineer Full-time Permanent EU

Job details

Not specified Salary
EU Eligibility
Not specified Experience
Full-time Employment

About this role

Role overview

This position sits at the intersection of research and infrastructure engineering, focused on building the distributed training systems that power large-scale multimodal foundation models running across thousands of GPUs. The work spans hard PyTorch and CUDA engineering alongside distributed-systems design, with an emphasis on advanced parallelism, training stability, and high utilization across massive clusters. It is built for engineers who have already solved real foundation-model training problems rather than those still ramping into distributed systems.

Responsibilities

- Design, implement, and optimize distributed training systems that scale to thousands of GPUs while preserving throughput and stability - Research and integrate advanced parallelization strategies, including FSDP, tensor, pipeline, and expert parallelism - Build monitoring, visualization, and debugging tooling that makes large training runs observable and diagnosable end to end - Tune training stability, convergence behavior, and resource utilization across massive GPU fleets

First 90 days

- Days 1–30 — Immerse and diagnose the current training stack, identifying where stability and utilization break down at scale - Days 30–60 — Ship a parallelization or stability improvement that measurably helps a real production training run - Days 60–90 — Build the monitoring and reliability tooling that keeps multi-thousand-GPU runs healthy and efficient

Requirements

- Extensive hands-on experience with distributed PyTorch training and the parallelisms used in foundation-model development - Deep understanding of GPU cluster architecture, including networking and storage subsystems - Familiarity with communication libraries such as NCCL and MPI, paired with a track record of distributed-system optimization

Nice to have

- Strong Linux systems administration and scripting skills - Experience managing training runs across 100+ GPU deployments - Background with containerization, orchestration, and cloud infrastructure

Detected Oct 8, 2026
Last verified Oct 8, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight