Remote job
Senior Network Engineer – Deployment (APAC)
Job details
About this role
Role overview
Hands-on Senior Network Engineer role based in Singapore, building and bringing up the fabrics behind GPU clusters across Asia-Pacific. The position sits within a deployment organization that owns every step from hardware arrival on site through to customer acceptance, taking network designs into production and making sure every build is ready for customers.
Responsibilities
- Deploy, configure, and validate high-performance Ethernet and InfiniBand fabrics for large GPU clusters at regional sites. - Bring up front-end, storage, and management networks and WAN edge connectivity for new sites and expansions. - Work with carriers, colocation providers, and IXPs across APAC to deliver circuits, cross-connects, and peering. - Run network acceptance testing and performance validation ahead of customer handover. - Troubleshoot and resolve issues across routing, switching, optics, and cabling during build and early life. - Apply global reference designs and build standards, feeding field learnings back to network architecture. - Build and improve automation for provisioning, configuration, and validation. - Write clear runbooks and documentation so builds are repeatable across sites and regions.
Requirements
- 5+ years of network engineering experience in large-scale data centre, cloud, or HPC environments. - Background at a hyperscaler or major network vendor, or another large-scale network operation. - Strong hands-on knowledge of routing and switching: BGP, OSPF, EVPN/VXLAN, ECMP, and L2/L3 data centre design. - Experience deploying and operating cluster networks, ideally including InfiniBand or RoCE fabrics. - Familiarity with the APAC data centre and carrier landscape and working with regional providers. - Practical network automation skills (Python, Ansible, or similar). - Calm, methodical troubleshooting under deadline pressure on live builds. - Willingness to travel across Asia-Pacific.
Nice to have
- Networks built specifically for GPU clusters or AI training environments. - HPC topologies such as Fat Tree or Rail, and 400G/800G optics. - SONiC, Cumulus, or other open networking platforms. - WAN, subsea or long-haul fibre, peering, or DWDM experience. - Standing up infrastructure in a new region or market. - Industry certifications like CCNP/CCIE, JNCIP/JNCIE, or NVIDIA networking certifications. - Mandarin written and verbal communication skills.