Remote job
Sr. Devops Engineer
Job details
About this role
Role overview
A senior DevOps, cloud, and site reliability engineering role is open on a platform team that builds and operates infrastructure for AI SaaS and PaaS products. The engineer will shape delivery pipelines, harden cloud foundations, and elevate platform observability and resilience. The position is remote within the country and may include future on-call participation.
Responsibilities
- Build, evolve, and operate cloud infrastructure backing AI software and platform services. - Design and refine CI/CD pipelines across multiple ecosystems to enable reliable software delivery. - Automate infrastructure provisioning and lifecycle management using Terraform. - Manage Kubernetes clusters, including cloud-managed variants, and administer Linux environments end-to-end. - Support PostgreSQL deployments, including operational troubleshooting and performance investigation. - Define service-level objectives and produce monitoring, alerting, dashboards, and escalation paths that surface real issues quickly. - Collaborate with engineering teams to clarify requirements, translate complex problems into actionable tasks, and improve platform reliability. - Produce and maintain clear technical and operational documentation. - Participate in an on-call rotation if the team adopts one in the future.
Requirements
- 7+ years of experience in DevOps, cloud engineering, or site reliability engineering. - 5+ years working with SDLC methodologies and CI/CD technologies, with hands-on expertise in at least two CI/CD ecosystems. - 3+ years as a software engineer, preferably with a backend focus. - 5+ years of Linux administration paired with strong networking and operating systems knowledge. - 5+ years of deep cloud experience on AWS (highly preferred) and/or GCP. - 5+ years with Infrastructure as Code tools, primarily Terraform, and 5+ years managing Kubernetes including at least one cloud-managed service such as EKS, GKE, or AKS. - Strong PostgreSQL experience and proven cross-layer diagnostic skills across code, cloud, network, and OS. - 5+ years with observability tooling, including designing SLOs, alerts, dashboards, and escalation chains. - Demonstrated ability to gather requirements, decompose complex problems, and drive focused discussions. - Collaborative, resilient, and proactive style with strong documentation habits.
Nice to have
- AWS or GCP certifications. - Experience integrating MLOps practices into CI/CD pipelines, exposure to Nvidia NIM, or relevant Nvidia certifications. - Familiarity with Temporal workflows.
Benefits and work setup
- Remote work arrangement based in Colombia. - Opportunity to contribute to platform engineering for AI-driven enterprise products.