Remote job
Cluster & Systems Capacity Engineer
Job details
About this role
Role overview This role sits within a cloud operations organization that runs a large-scale global cloud storage platform serving hundreds of thousands of customers. The engineer drives planning, deployment, and optimization of storage clusters, compute, and network infrastructure to ensure the platform scales reliably, cost-efficiently, and ahead of demand. It is a high-impact position directly contributing to service availability, durability, performance, margin optimization, and long-term platform scalability.
Responsibilities - Develop and maintain short, medium, and long-term capacity demand and hardware deployment forecasts across storage, compute, and network domains - Build predictive models that translate business demand signals into infrastructure requirements using historical utilization, growth trends, product plans, and hardware lifecycle roadmaps - Partner with infrastructure, production, and network engineering teams to align capacity plans with system design and scaling initiatives - Develop and automate forecasting pipelines, simulation calculators, capacity dashboards, and related tooling to improve data quality and provide clear visibility into platform usage and cluster health - Monitor and analyze cluster and system-level utilization and performance across CPU, memory, IOPS, and network resources - Adjust deployment plans and recommended configurations in real time to maintain adequate headroom and system stability - Partner with service and platform owners to develop headroom policies, optimize hardware bills of materials, leverage virtualized orchestration, and reduce product cost - Work with operations and finance peers to align capacity plans with capital budgets, cost targets, and financial outcomes - Lead evaluation, procurement, and provisioning of new or additional hardware in partnership with systems and network engineering, SRE, NOC, and data center operations teams
Requirements - Bachelor's degree in Computer Science, Engineering, Mathematics, Data Science, Information Systems, Statistics, or a related technical field, or equivalent experience - Three to six or more years of experience in site reliability engineering, infrastructure capacity planning, systems or infrastructure engineering, production engineering, data center operations, or a similar cloud operations role - Familiarity with cloud storage infrastructure, particularly highly available, large-scale distributed systems that support large data volumes with high throughput and complex performance requirements - Background in capacity modeling, performance analysis, scenario modeling, and infrastructure cost optimization, with ability to quantify ideas within financial frameworks - Proficiency with database and data analysis tools such as Snowflake, Metabase, Grafana, Python, SQL, Prometheus, Victoria Metrics, and Excel or Google Sheets - Strong data analysis, analytical, and logical reasoning abilities - Excellent communication and documentation skills, with the ability to explain complex concepts accurately and concisely - Desire to work on a highly autonomous team that cares deeply about quality, cost, and the customer experience
Benefits and work setup - Expected base salary range of $123,000 to $175,000, determined by factors including location, skills, experience, and relevant credentials - Healthcare coverage for family, including dental and vision - Competitive compensation plus a 401K plan - RSU grants for full-time employees and an employee stock purchase plan - Flexible vacation policy with maternity and paternity leave - MacBook Pro provided for work plus a generous workstation personalization stipend - Childcare bonus, fertility treatment and support benefits, commuter benefits, and a learning and development program - Culture that supports a healthy work-life balance