Remote job
Software Engineer, Infrastructure
Job details
About this role
Role overview A small platform team is hiring an infrastructure engineer to own the foundation that the rest of engineering runs on. The scope spans a major cloud platform, Kubernetes, a GPU fleet, CI/CD, monorepo health, and the security and cost boundaries around model training and inference. Expect deep collaboration with every other engineering team, real architectural influence on a small team with a large surface, and shared on-call responsibility.
Responsibilities - Operate and evolve the core platform: cloud infrastructure, Kubernetes, workflow orchestration, GPU capacity, and the deploy and rollback pipeline that ships every change. - Build the AI enablement substrate covering GPU capacity, training and inference pipelines, and the reliability and cost of model serving in production. - Treat cost as an engineering constraint, including the metering and attribution behind inference usage as it scales. - Implement security boundaries: identity and access, secrets management, least-privilege, and supply-chain integrity. - Produce infrastructure-as-code, runbooks, and observability that make systems legible to whoever joins next. - Improve how the team learns and ships through incremental releases, instrumentation, honest measurement, and clear technical writing and review.
Requirements - 8+ years building and operating production distributed systems, or equivalent server-side engineering with a heavy infrastructure focus. - Practical experience using AI and agent tooling to multiply your own impact, and a clear view of when to apply it. - A track record of running systems where failure was expensive, including carrying a pager and commanding incidents under pressure. - Daily use of SLOs and error budgets as operating tools. - Production experience with a major cloud provider and Kubernetes, with infrastructure-as-code as the default mode of work. - Owned an architecture or migration end to end whose consequences outlived the project, with lessons you can articulate. - History of finding unowned work, scoping it, earning buy-in, and shipping without a handed spec. - A debugging style built around minimal repros, careful log reading, and targeted checks rather than trusting output.
Nice to have - GPU or ML infrastructure
experience: capacity planning, training or inference pipelines, and serving cost and latency. - Production security engineering across IAM, secrets, and supply chain. - Cloud cost modeling, commitment strategy, reservations, and unit economics. - CI/CD at monorepo scale and developer environment work. - Background in media, video, or other GPU-backed workloads. - Experience on small platform teams covering a large surface area.
Benefits and work setup - Base salary range of $220,000 to $292,000, plus equity and benefits; final offer varies with experience and location. - Healthcare package, 401k matching, catered lunches, and flexible vacation time. - A mix of remote and hybrid roles, with headquarters in San Francisco and periodic in-person collaboration opportunities for remote teammates. - Equal opportunity employer committed to building a diverse team across backgrounds, experiences, and perspectives.