Remote job
Infrastructure Operations Engineer
Job details
About this role
Role overview An Infrastructure Operations Engineer is needed to help scale and operate a next-generation AI infrastructure platform that supports training and inference workloads at large scale. The InfraOps team sits at the center of reliability, automation, and operational scale for GPU infrastructure, owning break/fix operations, incident response, customer provisioning, observability, and the automation systems that keep complex infrastructure running efficiently. The role is hands-on, working across large-scale GPU environments, Linux systems, bare metal infrastructure, provisioning workflows, and platform reliability.
Responsibilities - Design, build, and roll out new platforms and patterns that minimize incidents and enable customer-facing and internal features - Deploy updates and improvements supporting both internal and end-customer use cases - Partner with Infrastructure Engineering, Network Operations, Customer Success, and Software/Platform Development teams to troubleshoot issues and improve operational efficiency - Participate in an evenly distributed on-call rotation following a primary/secondary pattern - Build automation that reduces manual toil over time across the infrastructure stack
Requirements - 8+ years working with Linux as a server/hosting platform, with Ubuntu experience a plus - 5+ years of experience with AWS - 2+ years of experience with Kubernetes and strong container fundamentals - 2+ years of experience with Terraform and Ansible - 2+ years managing network-attached storage via NFS, ceph, or similar protocols - Hands-on experience with GPU servers in bare metal or virtualized form - Deep experience with network switches, routers, and firewalls - Software development experience using Python, Go, bash, or similar for automation and integrating systems and APIs - Familiarity with monitoring systems such as Prometheus and the ELK stack, plus gitops workflows
Nice to have - Experience with VAST storage systems - Dell hardware troubleshooting and provisioning experience - Datacenter-level networking, 400Gb ethernet, and Infiniband experience - Experience with SONiC switches, Palo Alto firewalls, or Juniper Networks
Benefits and work setup - Anticipated annual base salary range of $160,000–$200,000 USD, plus discretionary bonus and meaningful equity - Comprehensive medical, dental, and vision coverage for employees and eligible dependents - 401(k) matching - Unlimited PTO, company holidays, and floating holidays - Two-week company-wide winter closure - Paid parental and family leave - Annual learning and development allowance - Wellness and work-from-home stipends - Four weeks of paid sabbatical leave after four years of service - Flexible schedules with a hybrid model for office-based teams - Fully remote within the U.S., or hybrid from NYC, SF, Seattle, or London, with occasional team and company offsites - Visa sponsorship is not available for this role