Remote job
Site Reliability Engineer
Job details
About this role
Role overview Join a Platform Engineering team supporting a cloud-native fraud and financial crime detection platform that processes billions of transactions annually. As a Site Reliability Engineer, you will focus on performance, reliability, and automation for distributed services running at ultra-low latency and high throughput.
Responsibilities - Provide capacity and scaling recommendations balancing cost, resilience, and performance. - Build automation, tooling, and infrastructure to eliminate manual operational work. - Develop and maintain infrastructure-as-code across the full lifecycle of production environments. - Participate in incident response, root cause analysis, and blameless postmortems. - Create actionable alerts and operational playbooks for the on-call rotation. - Partner with product teams to improve reliability and performance before and after release.
Requirements - Bachelor's degree in Computer Science, Information Systems, or equivalent practical experience. - Programming ability in Go, Python, or a comparable language. - 2+ years of experience with data structures, algorithms, and asynchronous or multithreaded systems. - 2+ years building scalable, distributed cloud services. - 2+ years operating production environments, including on-call responsibilities. - Strong written and verbal communication skills with a systematic approach to problem solving.
Nice to have - Hands-on experience with Kubernetes, AWS, GCP, or HashiCorp tooling. - Familiarity with observability stacks such as Grafana and Prometheus. - Experience collaborating across teams in a supportive, advisory engineering role.