← Back to jobs

Remote job

HPC/GPU Systems Engineer

Other Full-time Permanent US (Houston/SF/Seattle)

Job details

$100,000—$140,000 Salary
US (Houston/SF/Seattle) Eligibility
Senior Experience
Full-time Employment

About this role

Role overview This position sits within a Support and Operations team responsible for the day-to-day health of large-scale GPU fleets powering AI infrastructure. It is a hands-on L2/L3 engineering role covering tickets, alerts, hardware faults, GPU nodes, high-performance networking, Linux, and data centre operations, with a clear pathway toward senior scope through exposure to advanced AI environments.

Responsibilities - Own tickets and tasks end-to-end, escalating early when issues exceed scope and producing clean, evidence-rich handovers. - Perform GPU node triage and hardware troubleshooting, interpreting nvidia-smi or DCGM output and system logs, isolating faults across GPU, NIC, and server hardware, and carrying out physical remediation such as reseats and swap testing. - Run fabric and link diagnostics following established runbooks using tools like mlxlink, capture evidence accurately, and hand off cleanly to senior engineering when needed. - Assist with storage and data-path investigations on high-performance platforms, gathering evidence to support deeper diagnosis. - Maintain accurate records across DCIM, inventory, and asset systems, and contribute simple scripts and tooling improvements that reduce manual work. - Serve as the escalation point for onsite data centre operations staff, coordinating smart-hands tasks within scope.

Requirements - 3-4+ years in infrastructure support or support engineering roles within structured, customer-facing environments. - Working knowledge of GPU infrastructure and hands-on hardware troubleshooting experience. - Strong grasp of Linux fundamentals and the ability to read and interpret system logs. - Clear, structured written and verbal communication, treating documentation quality as a core engineering skill. - Discipline and organisation around ticket hygiene, follow-through, and runbook execution.

Nice to have - Familiarity with Kubernetes concepts such as nodes, pods, services, and logs. - Experience with Ansible, Terraform, CI/CD pipelines such as GitHub Actions, or access tooling such as Teleport or Vault. - Progress toward Linux, networking, Kubernetes, cloud, or security certifications.

Benefits and work setup - Remote-first working model with flexible scheduling. - Compensation includes base plus equity, with reviews every 12 months. - Listed base salary range is $100,000–$140,000 USD, with potential bonus, equity, or commission eligibility. - Medical, dental, vision, paid time off, parental leave, and retirement plan participation may be offered.

Skills detected in the listing

PythonKubernetesTerraform
Detected Sep 11, 2026
Last verified Sep 11, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight