Remote job
Technical Support Engineer (Bare Metal)
Job details
About this role
Role overview
A customer-facing technical support role focused on maintaining the reliability, performance, and scalability of large bare-metal GPU fleets that power AI workloads across multiple data centers. The position sits at the intersection of customer engineering, hardware operations, and infrastructure reliability, with opportunities to influence product and engineering direction through field feedback.
Responsibilities
- Deliver high-level technical support for customers running on bare-metal GPU cloud infrastructure, diagnosing and triaging reported issues and high-priority incidents across firmware, drivers, and hardware layers. - Develop a deep understanding of customer workloads and use cases to provide tailored troubleshooting and guidance. - Coordinate remote troubleshooting and physical hardware interventions with on-site data center technicians. - Create and maintain internal documentation, including troubleshooting guides, best-practice articles, and knowledge-base entries. - Participate in an on-call rotation to support production GPU clusters and ensure operational reliability. - Build automation and scripts (Python, Bash, Ansible, or similar) to streamline repetitive support workflows. - Partner with networking, hardware, and software engineering teams to resolve complex issues and feed recurring problems back into product improvements.
Requirements
- Hands-on experience in data centers, GPU clusters, large-scale server deployments, system administration, or hardware troubleshooting. - Intermediate proficiency with Linux (Ubuntu, CentOS, or similar) at the command line. - Working knowledge of GPU systems from vendors such as NVIDIA, server platforms such as SuperMicro and Dell, and high-performance computing environments. - Solid grasp of networking fundamentals (TCP/IP, VLANs, DNS, DHCP) and standard troubleshooting tools. - Experience with firmware updates, BIOS configuration, driver management, and multi-layer log analysis. - Familiarity with issue-tracking and documentation platforms (e.g., Jira, Confluence, Notion) plus scripting/automation experience.
Nice to have
- Curiosity about Kubernetes, Docker, and other containerized infrastructure technologies. - Comfort working across cross-functional teams and driving continuous improvement in fast-changing environments.
Benefits and work setup
- Medical, dental, and vision insurance fully covered for employees. - Company-paid life insurance plus short- and long-term disability coverage, flexible spending and health savings accounts. - 401(k) with employer match, employee stock purchase program eligibility, and tuition reimbursement. - Mental wellness benefits, family-forming support, paid parental leave, and full-service childcare assistance. - Flexible paid time off and a casual, innovation-focused work culture with catered meals at office and data center locations. - This position involves access to export-controlled information and is limited to U.S. persons as defined by applicable export regulations.