Remote job
Platform Engineer, AI/ML Infrastructure
Job details
About this role
Role overview A platform engineering role focused on the infrastructure that underpins AI/ML systems for organizations that want to keep ownership of their models, data, and results. The work covers the full stack of an AI platform: Kubernetes clusters tuned for GPU workloads, the platform services that turn clusters into usable systems, the delivery pipelines that ship releases reliably, and the operational practices that keep everything resilient. A defining constraint is that much of this must run in restricted or disconnected environments, so portability, reproducibility, and operability are design inputs from the first commit rather than problems handed downstream.
Responsibilities - Build and operate Kubernetes-based infrastructure for demanding AI/ML workloads, including GPU scheduling, resource management, and multi-tenant isolation. - Design and implement platform services for workflow orchestration, data ingest, model serving, results management, policy enforcement, and audit logging behind documented APIs. - Write infrastructure as code and build GitOps pipelines so environments are reproducible from source. - Build and operate CI/CD pipelines that produce versioned, signed, scanned release artifacts along with the documentation needed to deploy them. - Own reliability through capacity planning, upgrade paths, failure-mode analysis, backup and recovery, incident response, and postmortems. - Implement monitoring, logging, tracing, and alerting, and define the service level objectives they are measured against. - Deploy and validate the platform in restricted, disconnected, or limited-connectivity environments and verify parity after each release. - Write runbooks and operational documentation that other engineers can execute independently. - Contribute improvements back to upstream open-source infrastructure, Kubernetes, and MLOps projects.
Requirements - Hands-on experience operating Kubernetes at scale in production environments. - Familiarity with GPU scheduling, resource management, or other specialized compute concerns for AI/ML workloads. - Proficiency with infrastructure-as-code tooling such as Terraform or OpenTofu, and with GitOps workflows. - Experience building CI/CD pipelines with an emphasis on reproducible, signed, and scanned releases. - Strong understanding of observability practices, including monitoring, logging, tracing, and SLOs. - Eligibility to work in the United States and willingness to obtain and maintain a U.S. security clearance.
Nice to have - Experience supporting rapid prototyping programs or defense-adjacent innovation initiatives. - Familiarity with data sovereignty and privacy requirements for enterprise or government AI systems. - Background leading technical initiatives, setting engineering standards, or mentoring other engineers.
Benefits and work setup - Salary range of $120,000–$250,000 USD, dependent on experience and location. - Two openings: one U.S.-remote role and one hybrid role based in Washington, DC; Denver, CO; or Colorado Springs, CO, with up to 15% travel. - The hybrid role requires an active TS/SCI clearance; the remote role does not require an active clearance but candidates must be eligible to obtain and maintain one. - 100% employer-paid medical premiums for employees and self-managed PTO with a minimum time-off requirement.