Remote job
(EG0069) Senior Site Reliability Engineer (SRE) - Cassandra & AWS - Talent Connection
Job details
About this role
Role overview A senior SRE seat focused on modernizing a production customer-data caching platform that serves 20-25+ consuming applications at roughly 3,000 requests per second. The role works across SRE, DevOps, engineering, and architecture to scale the platform toward significantly more traffic, improve resiliency and automation, and shape deployment practices. It suits someone who works independently, sets technical direction, and thrives in small engineering teams.
Responsibilities - Design, build, and operate highly available, fault-tolerant distributed systems and caching platforms. - Evaluate current architecture and identify opportunities to scale toward 6x traffic while preserving performance and resiliency. - Improve caching strategies including TTLs, refresh patterns, cache placement, and downstream dependency management. - Move the platform toward active-active resiliency, validating failure, recovery, and capacity scenarios. - Build automated CI/CD pipelines with rolling, blue/green, or canary strategies and quality gates. - Establish infrastructure, configuration, secrets, and deployment practices through an everything-as-code approach. - Develop monitoring, alerting, and observability for cache latency, throughput, availability, and error rates, while defining SLIs, SLOs, and error budgets. - Lead root-cause analysis and blameless post-incident reviews to drive reliability improvements.
Requirements - Bachelor's degree in Computer Science, Engineering, or a related field. - 5+ years in SRE, DevOps, platform engineering, or a similar production infrastructure role. - Apache Cassandra experience covering data modeling, replication, consistency, tuning, and multi-datacenter deployments. - Hands-on AWS experience and familiarity with cloud-native design. - Strong observability and production troubleshooting skills, plus experience with performance testing and capacity planning. - Programming or scripting ability for automation and tooling, with Java/Spring Boot experience strongly preferred. - Ability to work independently, own ambiguous technical problems, and provide technical leadership within a small team. - Advanced English for working with U.S. clients.
Nice to have - GraphQL and API gateway experience. - Kubernetes and containerized application experience. - GitLab CI/CD and Terraform or similar infrastructure-as-code tools. - Prometheus, Grafana, Datadog, or comparable observability platforms. - Kafka, Amazon MSK, or other event-streaming technologies. - Active-active or multi-region architecture design experience.
Benefits and work setup - Full remote work within a LATAM team, with coworking spaces available across the region. - Competitive USD salary, paid time off at full salary per local rules, national holidays, sick leave, a yearly refundable health and well-being credit, and a birthday day off.