Remote job
Sr. Manager, Capacity Engineering
Job details
About this role
Role overview
This senior leadership role owns the capacity engineering function for a large-scale cloud platform, ensuring compute and GPU resources are forecasted, supplied, and managed reliably and cost-effectively. The position partners across infrastructure, SRE, finance, and product teams, and applies AI-assisted analysis to forecasting and operational workflows.
Responsibilities
- Set the 12–18 month strategy, roadmap, and success measures for the capacity engineering function - Develop CPU and GPU forecasts and supply plans that balance workload demand, delivery constraints, cost, and reliability - Guide the design of capacity requests, reservations, entitlements, allocation policy, and infrastructure data systems - Improve utilization and efficiency across Kubernetes-based shared compute platforms and GPU infrastructure while protecting availability and developer velocity - Establish trusted metrics and operating practices for capacity risk, allocation, infrastructure cost, realized savings, and system health - Lead cross-functional programs and make capacity and infrastructure-investment decisions with infrastructure, SRE, finance, product, and cloud provider stakeholders - Plan team structure and longer-term headcount, recruit senior talent, and develop technical leaders - Use AI tools to accelerate capacity analysis, forecasting, and operational investigations while applying judgment and verification to ensure quality
Requirements
- Bachelor's degree in computer science, a related field, or equivalent experience - Experience managing teams that build or operate large-scale cloud infrastructure, distributed systems, compute platforms, or resource management systems - Background defining 12–18 month technical strategy tied to organizational goals and delivering results through senior technical leaders - Hands-on familiarity with public cloud infrastructure, highly available distributed systems, Kubernetes, and capacity management systems - Experience leading cross-functional infrastructure programs and communicating availability, cost, and capacity tradeoffs to engineering, finance, SRE, and cloud provider stakeholders - Demonstrated ability to use AI to improve speed and quality in day-to-day work - Strong track record of critical evaluation and verification of AI-assisted outputs through testing, source-checking, and peer review
Nice to have
- History of building inclusive, accountable, and high-performing engineering organizations - Comfort with longer-term headcount planning and senior technical recruiting
Benefits and work setup
- Remote-eligible within the U.S. with in-person collaboration expected roughly 1–2 times per quarter - Relocation assistance is not provided - Annualized base compensation range of $208,592–$429,454 plus equity eligibility