Remote job
Senior DevOps Engineer - Storage Platforms
Job details
About this role
Role overview A senior DevOps position centered on storage engineering for a private cloud, building and automating large-scale Software Defined Storage and Kubernetes platforms. The role combines Infrastructure-as-Code, GitOps, and SRE practices to deliver secure, multi-tenant-isolated storage services that underpin high-throughput production workloads.
Responsibilities - Deploy, automate, and operate large-scale Software Defined Storage architectures across private and public cloud regions following ITIL practices. - Deploy and support enterprise storage platforms such as Pure Storage, HPE, and NetApp alongside SDS solutions like Ceph and Longhorn. - Integrate self-service storage workflows for Kubernetes CSI drivers and OpenStack consumers including VMs and bare metal. - Implement and manage backup solutions, with a preference for Rubrik-based platforms. - Build and maintain Infrastructure-as-Code for storage using Ansible, Terraform, Helm, and Git, with Python and Bash automation. - Build CI/CD pipelines for infrastructure updates including patching, upgrades, testing, and rollback. - Improve monitoring, alerting, and observability for storage capacity, latency, IOPS, and recovery health using GitOps and tools such as Prometheus, Loki, and Grafana. - Perform deep troubleshooting across storage, Kubernetes, hypervisors, networking, and Linux. - Author technical documentation, architecture diagrams, runbooks, and operational procedures. - Participate in on-call rotations, incident response, and root cause analysis, collaborating globally on change management.
Requirements - Six or more years managing enterprise storage and Kubernetes platforms on Linux. - Strong hands-on experience with SDS solutions such as Ceph and Longhorn, including storage migrations from legacy systems. - Expertise with block, file, and object storage, Fibre Channel (Cisco MDS), and IP-based protocols like NVMe-oF or iSCSI fabrics. - Expert-level Kubernetes and Linux (Ubuntu, RHEL, CentOS) operations skills. - Proficiency with Infrastructure-as-Code (Ansible, Terraform) for provisioning storage and backup schedules. - Backup technology experience, preferably Rubrik. - Strong scripting in Python and Bash; Go is a plus. - Experience operating 24x7 mission-critical production environments. - Hands-on experience with KVM hypervisors such as SUSE Harvester and OpenStack. - Strong written and verbal communication skills; fluency with Git, CI/CD pipelines, and automated testing. - Bachelor's degree in computer science or equivalent professional experience.
Nice to have - OpenStack Cinder multi-backend administration. - Familiarity with CIS/NIST security and infrastructure lifecycle management. - ITIL Foundation or advanced certifications. - Background in telco, edge cloud, or large enterprise environments. - Certifications such as CKA, CKS, or Red Hat Ceph Storage Administrator (EX125).
Benefits and work setup - Fully onsite, five days a week in office. - Compensation package includes performance bonus eligibility and equity, health, dental, and vision coverage from day one, mental health support resources, 401k matching, employee stock purchase plan, paid holidays, volunteer time, and 12 weeks of paid parental leave.