Remote job
Site Reliability Engineer
Job details
About this role
Role overview Design and operate the cloud platform underpinning a carrier-as-a-service telecom product that lets organizations spin up and scale their own mobile networks. The role combines infrastructure engineering with reliability practices: building automation, monitoring mission-critical systems, and enabling product, telecom, and data engineering teams to ship and operate their services confidently. Expect participation in an on-call rotation and a culture of continuous improvement through blameless postmortems.
Responsibilities - Design and implement cloud platform components that support backend telecom services. - Automate technical operations including deployments, scaling, and recovery workflows. - Monitor and maintain mission-critical production infrastructure to maximize uptime. - Participate in an on-call rotation and contribute to a culture of blameless postmortems and continuous improvement. - Enable engineering, telecom, and data engineering teams by providing the tools needed to operate the services they build.
Requirements - Solid understanding of Linux/Unix systems, including process management, filesystems, memory management, and networking. - Proficiency in at least one programming language such as Python, Go, or Ruby, plus strong scripting skills in Bash or Perl. - Hands-on experience with infrastructure provisioning tools such as Terraform, CloudFormation, or Ansible. - Familiarity with containerization using Docker and orchestration with Kubernetes. - Experience with monitoring and observability tools like Prometheus, Grafana, or Datadog, including alerting, log analysis, and dashboarding. - Experience participating in on-call rotations, handling incidents, and following practices such as runbooks and postmortems. - Experience building and maintaining CI/CD pipelines using tools such as Jenkins, GitLab CI, or CircleCI. - Hands-on experience with at least one major cloud provider such as AWS, Google Cloud, or Azure. - Understanding of TCP/IP, DNS, HTTP/HTTPS, load balancing, and firewalls, plus working knowledge of virtualization and cloud-native architecture.
Nice to have - Strong understanding of deployment strategies such as canary releases and blue-green deployments. - Familiarity with high availability design, failover mechanisms, IAM, and zero trust principles. - Experience with distributed systems such as Kafka, Cassandra, or Elasticsearch. - Functional knowledge of SQL and NoSQL databases. - Familiarity with distributed tracing tools like Jaeger or OpenTelemetry and log aggregation platforms such as the ELK stack or Splunk. - Experience with load testing, performance profiling, and configuration management tools like SaltStack.