Remote job
Senior Cloud Engineer
Job details
About this role
Role overview A Senior Cloud Engineer role focused on operating, monitoring, and continuously improving client cloud environments within a managed services context. The position leads incident response, drives automation, and supports 24x7 coverage through a rotating shift and on-call schedule. It is based in Brazil with remote or hybrid options in São Paulo or Florianópolis.
Responsibilities - Monitor cloud infrastructure across AWS, Azure, and/or GCP using observability tools, triaging alerts and responding to incidents within defined SLAs as part of a 24x7 rotation - Lead incident response, root cause analysis, and problem management, escalating to vendors or engineering leadership when required - Execute change requests, patching, backups, and routine maintenance following change management and least-privilege access processes - Build and maintain Infrastructure as Code using Terraform, CloudFormation, or ARM/Bicep, alongside CI/CD pipelines to automate provisioning - Improve monitoring, alerting, and dashboarding using tools like CloudWatch, Datadog, Grafana, or Azure Monitor to reduce mean time to detect and resolve - Support security operations by applying patches, responding to vulnerability findings, and following incident escalation procedures - Maintain runbooks, knowledge base articles, and shift handover notes for smooth continuity across teams and time zones - Collaborate with client stakeholders and internal engineering teams on service improvement initiatives - Contribute to capacity planning, cost optimization, and reliability improvements using SLOs and SLIs
Requirements - Advanced or fluent English - 5+ years of experience in cloud engineering, DevOps, SRE, or infrastructure support, ideally in managed services or 24x7 operations - Deep hands-on experience with at least one major cloud platform (AWS, Azure, or GCP), with multi-cloud experience preferred - Working knowledge of Linux and/or Windows Server administration - Scripting and automation experience with Python, Bash, or PowerShell - Solid experience with containerization and orchestration, specifically Docker and Kubernetes - Familiarity with monitoring and incident management tools such as Datadog, PagerDuty, ServiceNow, or Jira Service Management - Understanding of ITIL-aligned incident, problem, and change management practices - Willingness to work rotating shifts, including nights, weekends, and public holidays
Nice to have - AI-first approach to work, with strong analytical judgment for assessing AI outputs and systems thinking across incident-to-resolution chains - Experience with IT Service Management platforms enriched by AI
Benefits and work setup - Full-time position based in Brazil, with remote or hybrid options in São Paulo or Florianópolis - Remote and hybrid flexibility depending on country - Career advancement with international mobility and professional development programs - Access to learning resources, training, and industry experts - Inclusive culture with reasonable accommodations available during the interview process