Remote job
Site Reliability Engineer III (DBA)
Job details
About this role
Role overview A senior hands-on SRE position with primary ownership of production database systems, especially distributed MySQL via Vitess and Cassandra, within a large-scale cloud platform. The role helps establish the operational foundation of a new database-focused SRE function by designing architecture, writing runbooks, and serving as a senior escalation during incidents. It suits an engineer who combines deep database expertise with broad SRE skills across Linux, Kubernetes, observability, and automation.
Responsibilities - Design, deploy, and own highly available database architecture for Vitess (distributed MySQL) and Cassandra, partnering with data infrastructure teams on resharding and capacity planning. - Create operational procedures, runbooks, and escalation guidance used by more junior SRE database engineers. - Tune performance through query optimization, indexing, schema design, and capacity modeling, and own backup, recovery, replication, and disaster recovery testing. - Participate in on-call rotations and incident response, serving as an escalation point for complex database incidents and leading root cause analysis and post-incident reviews. - Build automation that reduces operational toil, and contribute to monitoring, logging, and alerting frameworks integrated with incident workflows. - Drive database security, access control, patching, hardening, and compliance practices in collaboration with engineering and security stakeholders.
Requirements - 6 to 8 years in SRE, systems engineering, infrastructure operations, or database engineering, with meaningful production database experience. - Deep hands-on experience with MySQL and distributed or sharded database systems; Vitess experience strongly preferred. - Experience administering NoSQL systems such as Cassandra and designing HA topology, replication, backup, and disaster recovery. - Strong Linux administration and SQL skills, including query performance analysis, indexing, and schema design. - Solid grasp of monitoring, alerting, incident response, SLIs, SLOs, and error budgets. - Experience with Kubernetes and Docker, infrastructure-as-code and configuration management tools such as Terraform, Ansible, and Jenkins, and proficiency in Python, Bash, or Go.
Nice to have - Background in SaaS, cloud services, or large-scale distributed systems. - Cloud platform experience with AWS, GCP, or Azure. - Familiarity with ITIL or OSS practices and SLA or SLO management. - Experience mentoring or onboarding engineers into complex technical environments.
Benefits and work setup - US-based role with an expected base salary range of $125,000 to $150,000, with final offer varying by location, experience, and credentials. - Family healthcare including dental and vision, 401(k), RSU grants, ESPP, flexible vacation, parental leave, childcare and fertility support, commuter benefits, and a learning and development program. - Company laptop plus a workstation personalization stipend and a culture that supports work-life balance.