Remote job
Staff Software Engineer - Fleet Management
Job details
About this role
Role overview This Staff Software Engineer role builds a fleet management workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale. The position sits at the intersection of distributed systems, infrastructure automation, and physical hardware, owning domain-level architecture in a Python-based system that manages the full operational lifecycle of GPU infrastructure.
Responsibilities - Own domain-level technical architecture for areas such as provisioning, validation, or remediation, influencing engineers across the team and adjacent squads - Design and build production-grade automation in Python for device provisioning, burn-in testing, network configuration, and hardware health validation - Engineer for reliability and auditability, treating idempotency, resumability, checkpointing, retries, replay, and failure handling as first-class design concerns - Integrate Fleet Manager with datacenter inventory systems, cloud orchestration platforms, and bare-metal provisioning tools - Operate what is built with strong observability, alerting, incident response, and day-2 operational discipline - Mentor and influence through design reviews, implementation guidance, and operational best practices
Requirements - Extensive experience designing, building, and operating distributed systems in production, ideally in infrastructure automation, workflow tooling, or platform engineering - Strong proficiency in Python and a deep understanding of event-driven and workflow architecture including reliable delivery, idempotency, retries, replay, and failure handling - Track record of taking automation systems from ambiguous requirements to production, including day-2 operations such as incident response and performance optimization - Proven ability to lead ambiguous technical work across team boundaries and drive domain-level delivery through influence rather than formal authority - Strong communication skills for building consensus with stakeholders in a fast-paced, high-agency environment
Nice to have - Experience with workflow orchestration tools such as Temporal, Airflow, or Prefect - Hands-on experience with infrastructure tooling like DCIMs, NetBox, OpenStack, MAAS, Ironic, IPMI, PXE boot, or network automation - Background in GPU infrastructure, HPC networking, or datacenter topology including high-performance interconnects - Deep knowledge of Kubernetes, Infrastructure as Code (Terraform or Pulumi), AWS, and GCP
Benefits and work setup - Base salary range of $220,000 to $320,000 USD, with potential eligibility for bonus, equity, and/or commission programs - Competitive benefits package that may include medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation