Remote job
Infrastructure Operations Engineer (APAC)
Job details
About this role
Role overview
Join a small operations team as a senior infrastructure engineer responsible for keeping a large-scale GPU and bare-metal compute fleet healthy. The work spans the full infrastructure stack—Linux hosts, compute, networking, storage, provisioning, and cluster orchestration—and centers on being the technical escalation point when complex issues arise. It is a hands-on, generalist role with a strong automation mindset rather than a traditional ticket-driven administration position.
Responsibilities
- Operate and troubleshoot production infrastructure across Linux servers, GPUs, networking, storage, and orchestration systems. - Serve as the escalation point for hard problems, driving incidents from initial investigation through root-cause analysis and resolution. - Own end-to-end operational workflows including provisioning, configuration, validation, maintenance, remediation, and decommissioning. - Build automation and internal tooling that removes repetitive work and lets the fleet scale without linear headcount growth. - Identify recurring failure modes and partner with infrastructure engineering teams to make systems more reliable and repeatable. - Participate in a primary/secondary on-call rotation supporting production infrastructure.
Requirements
- Significant hands-on experience operating and troubleshooting Linux-based infrastructure at scale. - Practical experience with bare-metal servers, GPUs, and the surrounding hardware and kernel stack. - Familiarity with hardware management and provisioning interfaces such as PXE, BMC, IPMI, Redfish, or iDRAC. - Experience with distributed or high-performance storage systems (e.g., VAST, Ceph, GPFS, or WEKA). - Experience with infrastructure-as-code, configuration management, or GitOps workflows. - Residing in Singapore and able to work Monday–Friday, 8:00 AM–5:00 PM local time (UTC+8) with periodic on-call.
Nice to have
- Background operating or troubleshooting high-speed data center networking. - Comfort switching between deep debugging of a single failure and stepping back to design systemic improvements.
Benefits and work setup
- Remote role based in Singapore with a Monday–Friday daytime window plus after-hours on-call participation. - Anticipated annual base salary range of SGD 165,000–220,000, plus discretionary bonus and meaningful equity. - Comprehensive medical, dental, and vision coverage for employees and eligible dependents. - Retirement savings support (401(k) matching or pension contributions depending on location). - Unlimited paid time off, company holidays, floating holidays, and a two-week company-wide winter break. - Paid parental and family leave, annual learning and research budget, and wellness plus work-from-home stipends. - Four-week paid sabbatical after four years of service, flexible scheduling, and in-office meals at company hubs.