Remote job
Forward Deployed Engineer - Physical AI
Job details
About this role
Role overview A senior, high-autonomy individual contributor position focused on owning the cloud infrastructure foundation that makes a physical AI platform fast, reliable, scalable, secure, and cost-effective. The role is embedded inside strategic customer and ISV engineering teams, shipping production infrastructure for real physical AI workloads — simulation, training, evaluation, inference, and batch — rather than demos. Remote work within the United States is welcome, with a preference for candidates in the SF Bay Area or Austin, TX.
Responsibilities - Own end-to-end technical execution inside each account: discovery, scoping, infrastructure design, build, and production rollout. - Build and operate cloud infrastructure powering customer physical AI workflows, owning compute orchestration for simulation, training, evaluation, inference, and batch workloads. - Build platform services for job execution, scheduling, retries, observability, logging, secrets, access control, and cost tracking. - Build onboarding infrastructure for pilots, including sandbox environments, dataset storage, workflow execution, and isolated, observable early deployments. - Optimize cost, utilization, performance, and reliability across workloads, and debug infrastructure issues across application, network, storage, compute, and orchestration layers. - Partner with other FDEs to expose infrastructure capabilities through clean APIs, SDKs, and product workflows. - Turn repeated customer infrastructure pain into reusable platform capabilities and partner with Product and Engineering to fold them into the core platform. - Use modern AI coding tools as primary leverage, treating engineering velocity as a primary success metric.
Requirements - Extensive experience building production cloud infrastructure for AI/ML workloads, including training pipelines, inference workloads, or HPC-style compute environments. - Proven ability to debug infrastructure issues across application, network, storage, compute, and orchestration layers. - Strong instincts for isolation, RBAC, uptime, and traceability on customer-facing workloads. - High agency with a bias toward simple, composable infrastructure that serves real customer workflows. - Strong written and verbal communication skills, comfortable in technical conversations with customer CTOs and able to debrief engagements to senior leadership.
Nice to have - Prior experience as a Forward Deployed Engineer or in an equivalent customer-embedded engineering function at a frontier company. - Experience with major AI cloud infrastructure providers. - Experience with Slurm, Soperator, Kubernetes GPU scheduling, Ray, Argo, Airflow, Metaflow, or similar orchestration tools. - Familiarity with NVIDIA GPU infrastructure, CUDA workloads, Isaac Sim, Omniverse, or simulation-at-scale.
Benefits and work setup - Competitive compensation with career growth opportunities. - Remote work within the United States. - Flexibility, ownership, and a collaborative, innovative culture. - Opportunity to contribute to impactful AI projects alongside an international team.