← Back to jobs

Remote job

Infrastructure Operations Engineer

DevOps Full-time Permanent UK and US

Job details

$160,000—$200,000 Salary
UK and US Eligibility
Lead Experience
Full-time Employment

About this role

Role overview An Infrastructure Operations Engineer is needed to help scale and operate a next-generation AI infrastructure platform that supports training and inference workloads at large scale. The InfraOps team sits at the center of reliability, automation, and operational scale for GPU infrastructure, owning break/fix operations, incident response, customer provisioning, observability, and the automation systems that keep complex infrastructure running efficiently. The role is hands-on, working across large-scale GPU environments, Linux systems, bare metal infrastructure, provisioning workflows, and platform reliability.

Responsibilities - Design, build, and roll out new platforms and patterns that minimize incidents and enable customer-facing and internal features - Deploy updates and improvements supporting both internal and end-customer use cases - Partner with Infrastructure Engineering, Network Operations, Customer Success, and Software/Platform Development teams to troubleshoot issues and improve operational efficiency - Participate in an evenly distributed on-call rotation following a primary/secondary pattern - Build automation that reduces manual toil over time across the infrastructure stack

Requirements - 8+ years working with Linux as a server/hosting platform, with Ubuntu experience a plus - 5+ years of experience with AWS - 2+ years of experience with Kubernetes and strong container fundamentals - 2+ years of experience with Terraform and Ansible - 2+ years managing network-attached storage via NFS, ceph, or similar protocols - Hands-on experience with GPU servers in bare metal or virtualized form - Deep experience with network switches, routers, and firewalls - Software development experience using Python, Go, bash, or similar for automation and integrating systems and APIs - Familiarity with monitoring systems such as Prometheus and the ELK stack, plus gitops workflows

Nice to have - Experience with VAST storage systems - Dell hardware troubleshooting and provisioning experience - Datacenter-level networking, 400Gb ethernet, and Infiniband experience - Experience with SONiC switches, Palo Alto firewalls, or Juniper Networks

Benefits and work setup - Anticipated annual base salary range of $160,000–$200,000 USD, plus discretionary bonus and meaningful equity - Comprehensive medical, dental, and vision coverage for employees and eligible dependents - 401(k) matching - Unlimited PTO, company holidays, and floating holidays - Two-week company-wide winter closure - Paid parental and family leave - Annual learning and development allowance - Wellness and work-from-home stipends - Four weeks of paid sabbatical leave after four years of service - Flexible schedules with a hybrid model for office-based teams - Fully remote within the U.S., or hybrid from NYC, SF, Seattle, or London, with occasional team and company offsites - Visa sponsorship is not available for this role

Skills detected in the listing

PythonGoAWSKubernetesTerraform
Detected Sep 9, 2026
Last verified Sep 10, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight