Remote job
Principal Software Engineer - Fleet Management
Job details
About this role
Role overview A Principal Software Engineer role serving as the technical lead for a fleet management platform that automates the lifecycle of GPU nodes and network switches at data center scale. The position combines architecture ownership, hands-on Python development, and mentorship of senior engineers building workflow orchestration for provisioning, burn-in testing, monitoring, and self-healing infrastructure.
Responsibilities - Lead technical architecture and roadmap for workflow automation systems that manage the full lifecycle of compute infrastructure. - Own end-to-end delivery of device provisioning, validation, testing, and remediation workflows at scale. - Design workflow orchestration systems covering GPU nodes and network switches, including enrolment, burn-in, configuration, health monitoring, and self-healing. - Establish engineering standards for reliability, observability, and operational excellence across all platform services. - Mentor senior engineers through design reviews, technical leadership, and hands-on collaboration. - Integrate with infrastructure tooling such as DCIMs, NetBox, OpenStack, and bare metal APIs including MAAS, Ironic, and IPMI. - Partner with infrastructure, platform, and SRE teams to translate operational needs into robust automation. - Build production-grade Python systems for hardware lifecycle automation, leveraging AI development tools to accelerate delivery.
Requirements - Twelve to fifteen or more years of software engineering experience building and operating production systems, with proven technical leadership in infrastructure automation or workflow tooling. - Strong Python engineering fundamentals and experience leading complex, multi-service distributed systems. - Deep understanding of operational excellence including SLOs, monitoring, alerting, incident response, and production reliability. - Track record of owning technical roadmaps and delivering large-scale automation from ambiguous requirements to production. - Regular use of AI coding tools as a core part of the development workflow. - Strong mentorship and stakeholder communication skills.
Nice to have - Hands-on experience with workflow orchestration tools such as Temporal, Airflow, or Prefect. - Background in bare metal provisioning and network automation, including PXE boot. - GPU infrastructure experience covering health monitoring, burn-in testing, or cluster management. - Familiarity with HPC networking including InfiniBand or RoCE, and with Kubernetes, Terraform or Pulumi, AWS, and GCP. - Open-source contributions in infrastructure automation or cloud-native tooling.
Benefits and work setup - Base salary range of $240,000 to $400,000 USD, with potential eligibility for bonus, equity, or commission programs. - Benefits package may include medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.