Remote job
Site Reliability Engineer ll (DBA)
Job details
About this role
Role overview A mid-level SRE role with a database administration focus, supporting production database systems built primarily on Vitess (distributed MySQL) and Cassandra. The position executes against architecture and runbooks established by senior DBA SREs, while building automation, maintaining observability, and contributing to incident response. It suits an engineer with solid fundamentals who is ready to grow in a large-scale, Kubernetes-based environment.
Responsibilities - Operate and maintain high-availability Vitess and Cassandra databases against established runbooks, escalating architectural changes to senior DBA SREs. - Optimize database performance through query tuning, indexing, and schema improvements, and execute documented backup, recovery, and replication procedures. - Support service health monitoring using SLIs, SLOs, and error budgets, and participate in on-call rotations, incident response, and post-incident reviews. - Develop automation for routine operational tasks and contribute to monitoring, logging, and alerting frameworks. - Work with CI/CD pipelines, configuration management, and infrastructure-as-code tools, and write scripts in Python, Bash, or Go to improve reliability. - Partner with engineering, product, and operations teams on capacity planning, disaster recovery exercises, and reliability-focused documentation.
Requirements - 2 to 4 years of experience in site reliability, systems engineering, or operations centered on database systems. - Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience. - Solid Linux systems administration and troubleshooting skills, with understanding of containers and microservices concepts. - Hands-on MySQL experience covering performance tuning, replication, and disaster recovery; direct Vitess or other sharded MySQL experience is a plus. - Proficiency in SQL and NoSQL database management, plus familiarity with monitoring, alerting, and root cause analysis. - Scripting ability in at least one of Python, Bash, or Go, and comfort operating in Kubernetes environments.
Nice to have - Experience in a SaaS, service provider, or large-scale distributed systems environment. - Familiarity with ITIL or OSS practices and SLA or SLO management. - Cloud platform experience with AWS, GCP, or Azure. - Ability to work independently and drive projects from problem discovery through resolution.
Benefits and work setup - Family healthcare including dental and vision, 401(k), RSU grants for full-time employees, ESPP, and a flexible vacation policy. - Parental leave, childcare and fertility support, commuter benefits, learning and development program, and a culture that supports healthy work-life balance. - Company laptop plus a workstation personalization stipend.