Remote job
Designated Service Engineer - Ceph Expert
Job details
About this role
Role overview
This is a senior, customer-facing engineering position focused on Ceph-based object storage for enterprise and research environments. The role blends deep hands-on Ceph architecture with a high-touch designated services approach, owning the design, deployment, lifecycle operations, and performance of large-scale storage clusters for strategic accounts. It also serves as the technical liaison between customers and internal engineering/product teams, driving reliability improvements and elevating the overall customer experience.
Responsibilities
- Architect, deploy, and operate large-scale production Ceph clusters exposing S3, balancing availability, performance, and operational simplicity. - Own cluster lifecycle activities including upgrades, patching, configuration management, routine health checks, and proactive risk remediation. - Lead incident response and root-cause analysis across the Ceph stack, hardware, OS, and networking layers. - Create and maintain runbooks, operational best practices, and customer-facing documentation to improve reliability, observability, and automation. - Serve as the primary technical liaison with internal engineering and product teams, tracking issues through the ticketing system and delivering clear, timely updates. - Proactively monitor customer environments, identify risks before they impact production, and advise on hardware and topology choices aligned to workload requirements. - Participate in on-call and follow-the-sun rotations, with occasional off-hours and travel as needed.
Requirements
- 10+ years in customer-facing technical roles solving complex enterprise infrastructure issues. - 5+ years of hands-on Ceph experience in production, covering cluster design, deployment, upgrades, and day-2 operations. - Strong understanding of Ceph internals, including MON quorum, MGR active/standby, OSD behavior, CRUSH rules and maps, pools, placement groups, and recovery/backfill dynamics. - Experience operating multi-petabyte Ceph environments and managing PG scaling, long-running recovery, and fleet-size challenges. - Practical experience with Ceph RGW and S3 concepts such as buckets, users/tenants, load balancing, scaling patterns, and performance troubleshooting. - Strong Linux/Unix administration skills in multi-platform, distributed environments, plus deep networking knowledge (Infiniband, Ethernet, DPDK, UCX). - Experience with observability stacks such as Prometheus and Grafana, and excellent written and verbal communication skills for both technical and non-technical audiences.
Nice to have
- Familiarity with Kubernetes, containers, or major cloud platforms (AWS, Azure, OCI, GCP) in storage-heavy deployments. - Experience supporting HPC or AI/ML infrastructure, including GPU clusters, high-throughput networking, and performance benchmarking. - Exposure to infrastructure-as-code or configuration management tools such as Ansible or Terraform. - Comfort collaborating across Support, Engineering, and Product using tools like Jira, Confluence, and Slack.
Benefits and work setup
- Customer-facing premium services role with exposure to cutting-edge storage technologies and top-tier enterprise and research organizations. - Opportunity to grow beyond Ceph into broader object-storage and adjacent designated services engagements. - Collaborative, accountable culture with structured onboarding, knowledge sharing through FAQs, KB articles, and reusable playbooks, and a clear path toward trusted technical leadership for assigned accounts.