Remote job
Expert Automation & Observability Engineer
Job details
About this role
Role overview
Serve as a senior technical architect for enterprise observability, application performance monitoring, and telemetry across hybrid cloud, Kubernetes, and legacy environments. The mission is to replace fragmented, reactive monitoring with a unified, automated, proactive operating model that improves reliability, reduces alert noise, and supports consistent service operations.
Responsibilities
- Design and govern observability architectures covering metrics, logs, traces, and events, including telemetry pipelines, retention, cardinality, and cost controls. - Lead first-of-a-kind implementations, turning evaluations of new technologies into secure, repeatable, production-ready patterns. - Define SLIs, SLOs, and error budgets, and act as a senior escalation point during major incidents and root-cause investigations. - Improve detection and recovery through event correlation, dynamic thresholds, dependency mapping, and automation that reduces MTTD and MTTR. - Drive observability-as-code and infrastructure automation using tools such as Ansible, Terraform, Python, Bash, and GitOps workflows. - Design monitoring for containers, Kubernetes, microservices, and multi-cloud environments, with secure telemetry using RBAC, TLS, secrets management, and image scanning. - Lead knowledge-transfer programs, vendor transitions, operational-readiness handovers, and mentoring for global 24x7 teams.
Requirements
- At least 12 years of total IT experience, including 5–7 years in a lead architect, SRE, or principal observability engineering capacity. - Extensive experience with observability, APM, telemetry, monitoring, and time-series platforms. - Strong experience across Kubernetes, Docker or OpenShift, public cloud environments, Linux, Windows Server, and enterprise infrastructure dependencies. - Hands-on capability with automation, infrastructure as code, CI/CD, ITSM integrations, and monitoring-agent or collector deployment. - Experience leading first-of-a-kind rollouts, complex operational transitions, and knowledge-transfer programs. - Advanced major-incident management and evidence-based root-cause analysis experience.
Nice to have
- Certifications such as Certified Kubernetes Administrator, a cloud architect credential, or an observability/APM vendor certification.
Benefits and work setup
- Primarily remote work is available when client-site presence is not required, with office-based options and flexible scheduling. - The source describes paid time off, health, dental, vision, disability, life, retirement matching, family-forming support, parental leave, education assistance, wellness resources, and sabbatical leave.