Remote job
DevOps Engineer, Cloud Infra
Job details
About this role
Role overview
A senior infrastructure role focused on operating and evolving the cloud and data-streaming backbone behind a large-scale, consumer-facing digital-asset platform. The position blends hands-on reliability engineering with internal platform development, partnering closely with product teams to keep global services performant and resilient. The team is also actively exploring how AI and LLM tooling can reduce operational toil and sharpen incident response.
Responsibilities
- Own production Kafka and Redis clusters end-to-end, including deployment, monitoring, scaling, and deep root-cause troubleshooting. - Respond to live incidents, lead post-incident reviews, and drive structural fixes that improve long-term stability. - Manage and optimize AWS and secondary cloud environments for cost, performance, and reliability. - Build internal DevOps platform capabilities such as online load testing and change management tooling. - Partner with application engineers to streamline releases and improve the deployment pipeline. - Integrate LLM and AI frameworks (for chat-based operations, intelligent alert triage, and root-cause analysis) into operational workflows.
Requirements
- At least five years operating Kafka and Redis in large-scale production environments, with the ability to collaborate with developers on code-level optimizations. - Proficiency in at least one of Python, Go, or Java, plus solid working SQL skills. - Hands-on experience with Docker and Kubernetes for container orchestration. - Strong background with CI/CD tooling such as GitHub Actions, Ansible, and Terraform. - Three or more years working on AWS; exposure to GCP, Azure, or Alibaba Cloud is considered a plus. - Demonstrated problem-solving ability and a collaborative, partnership-oriented working style.
Nice to have
- Practical experience building AIOps capabilities such as anomaly detection, alert correlation, automated remediation, or root-cause analysis. - Familiarity with LLM-based DevOps automation, including chat-ops assistants or AI-driven observability pipelines. - Hands-on work with frameworks such as Dify, Agno, or LangChain.
Benefits and work setup
- Remote-friendly arrangement, with possible adjustments depending on the nature of the team. - Results-oriented culture with autonomy over how work is approached, plus structured opportunities for career growth and continuous learning. - Competitive compensation and standard company benefits.