Remote job
Senior Database Administrator
Job details
About this role
Role overview A senior database administrator is needed to own production reliability for the data layer behind a high-volume online sportsbook, casino, and social gaming platform. The role centers on multi-region CockroachDB and PostgreSQL clusters that process bets, wagers, settlements, and payouts where every transaction must be fast and accurate. The position blends deep database internals work, incident response, tooling development, and AI-augmented engineering on a team that treats data infrastructure as a product.
Responsibilities - Operate and evolve distributed CockroachDB clusters and PostgreSQL instances across production, staging, and development environments. - Investigate contention, serialization retries, range hotspots, leaseholder placement issues, monotonic-key write pressure, and cross-region latency in transaction paths. - Trace slow queries through execution plans and optimizer behavior to the schema or call path that caused the regression, then codify safer rollout patterns. - Build database observability with Prometheus, Grafana, Mimir, Loki, and Snowflake, plus alerting on leading indicators and SLO tracking. - Join the on-call rota, lead incident analysis when the database is part of the failure, and drive blameless postmortems toward durable fixes. - Develop Go and Python tooling for log collection, explain-plan analysis, migration checks, capacity modeling, and runbook generation, and automate provisioning, backup, and recovery with Terraform and related infrastructure-as-code tooling. - Implement and audit access controls, encryption, audit logging, and retention practices required for a regulated gaming platform.
Requirements - 5+ years operating production relational databases as a DBA, Database Reliability Engineer, or Data Platform Engineer. - Deep production experience with PostgreSQL, CockroachDB, or another RDBMS, including fluency in isolation, consensus, locality, and query planning tradeoffs. - Strong understanding of SQL engine internals such as MVCC, the Volcano iterator model, and cost-based optimizer frameworks like Cascades, applied to indexing strategy and plan analysis. - Working knowledge of distributed-systems fundamentals including consensus protocols (Raft, Paxos), distributed transactions, consistency models, and failure modes. - Proficiency writing tools and automation in Python or Go, plus the ability to read and suggest fixes to line-of-business code in Java or a similar language. - UNIX/Linux administration skills with shell scripting, comfort on a cloud platform (AWS preferred), and willingness to participate in an on-call rotation.
Nice to have - Experience using AI engineering harnesses such as Claude Code or Codex for incident investigation, runbook drafting, and tooling, with disciplined verification against production evidence. - Exposure to AI-assisted observability and anomaly detection, and to evaluating what works in a regulated, money-moving environment. - Track record of building documentation, standards, and reusable patterns that other teams build against, and mentoring engineers on database and distributed-systems fundamentals.
Benefits and work setup - Hybrid/remote working environment with team gatherings across the US and Europe. - Employer-sponsored training and conference attendance. - Access to AI-first tooling and a startup-style culture backed by an established, security-conscious organization.