Remote job
Platform / Infrastructure Engineer
Job details
About this role
Role overview This role keeps a fleet of production AI agents running reliably across many client environments where downtime translates directly into lost revenue. The engineer will own the cloud platform, deployment tooling, and observability stack that lets the system scale while staying fast and resilient.
Responsibilities - Design and operate cloud infrastructure on AWS or GCP that hosts AI agents at scale - Build CI/CD pipelines that ship agent updates quickly and safely across multiple client tenants - Implement monitoring, alerting, and observability for agent health and runtime performance - Manage containerized workloads and orchestration using Docker and Kubernetes - Tune infrastructure spend while pursuing very high availability targets - Develop tooling for tenant isolation, configuration management, and secret handling
Requirements - At least three years of experience in infrastructure, DevOps, or platform engineering - Deep familiarity with AWS or GCP services and architecture patterns - Proficiency with infrastructure-as-code tools such as Terraform, Pulumi, or CloudFormation - Hands-on experience with Kubernetes or ECS for container orchestration - Solid understanding of networking, security fundamentals, and Linux systems - Willingness to participate in on-call rotations and lead incident response
Nice to have - Experience scaling machine learning or AI workloads, including GPU management and model serving - Familiarity with edge computing or fleets of IoT devices - Background designing multi-tenant SaaS architectures - Experience operating high-availability systems for mission-critical applications