Remote job
Infrastructure Engineer (Storage)
Job details
About this role
Role overview A platform company building developer tools and large-scale compute for AI/ML workloads is hiring a Storage Infrastructure Engineer to own the data plane of its distributed storage stack. The role sits at the intersection of software, hardware, and operations, focusing on high-throughput, low-latency data access for training, inference, and HPC workloads running on bare-metal infrastructure. It is a builder-focused position emphasizing automation, reliability, and scaling of production systems.
Responsibilities - Operate and scale distributed storage systems, including VAST and S3-compatible object storage such as Ceph - Troubleshoot complex storage and data-path issues spanning hardware, OS, and software layers - Build Python-based automation for provisioning, lifecycle management, and monitoring of storage clusters - Manage Linux-based production systems in bare-metal environments, partnering with data center teams on hardware bring-up, upgrades, and issue resolution - Support capacity planning, utilization tracking, and forecasting using monitoring and telemetry - Collaborate with infrastructure, network, and platform teams on design discussions, integration, and best practices for high-performance storage
Requirements - 5+ years in infrastructure engineering, systems engineering, or a closely related field - Hands-on experience operating distributed storage systems such as VAST, Ceph, or similar at scale - Strong Linux systems experience in production environments - Proficiency in Python or comparable scripting/programming languages for automation - Experience with bare-metal infrastructure and ability to debug issues across storage, OS, hardware, and networking boundaries - Familiarity with storage networking protocols (e.g., NFS) plus capacity planning, monitoring, and performance tuning
Nice to have - Production experience with VAST storage systems - Operating S3-compatible object storage at scale - Data center operations experience with physical hardware - Familiarity with AI/ML or HPC workloads and their storage requirements - Background in high-performance or low-latency distributed systems - Experience with RDMA, GPU Direct Storage, or supporting GPU-based workloads and large compute clusters
Benefits and work setup The anticipated base salary range is $180,000–$220,000 USD, with a discretionary bonus and meaningful equity component. The role may be fully remote within the U.S. or hybrid out of major office hubs, with occasional team and company offsites; visa sponsorship is not available for this position. The package includes medical, dental, and vision coverage; retirement matching; unlimited PTO plus company and floating holidays; a two-week company-wide winter break; paid parental and family leave; an annual learning and development allowance; wellness and work-from-home stipends; a four-week paid sabbatical after four years of service; flexible schedules; and complimentary in-office meals. Benefits vary by location, team, and role.