Remote job
Operational Data & Observability Engineer
Job details
About this role
Role overview A hands-on engineering role inside an AI-focused GPU cloud company, responsible for building and evolving the monitoring, logging, and observability stack that supports production infrastructure and applications. The position partners with DevOps, SRE, platform, and software engineering teams to deliver actionable operational data and improve reliability at scale.
Responsibilities - Design and implement observability strategies spanning infrastructure, services, and applications, including dashboards, alerts, and SLOs. - Build and maintain centralised logging and log analysis pipelines, and implement distributed tracing across microservices. - Establish performance baselines and develop anomaly detection strategies. - Deploy, configure, and maintain telemetry collection systems, and design the operational data pipelines that feed monitoring and analytics. - Develop APIs and integrations that make operational data consumable across teams, while managing data quality, retention, and storage costs. - Participate in on-call rotations, support incident response, and maintain runbooks, documentation, and troubleshooting guides. - Administer and enhance observability platforms such as Datadog, Grafana, Prometheus, the ELK Stack, New Relic, or equivalents, and automate their deployment and configuration.
Requirements - Three or more years of experience in DevOps, SRE, operations engineering, platform engineering, or observability engineering. - Hands-on experience with monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalents. - Experience with centralised logging platforms including the ELK Stack, Splunk, CloudWatch, or similar. - Proficiency with scripting or programming languages such as Python, Go, Bash, or equivalents. - Strong grasp of observability fundamentals covering metrics, logging, distributed tracing, and APM. - Experience with at least one major cloud platform and Kubernetes or other container orchestration technologies. - Solid understanding of performance monitoring across applications, infrastructure, networking, databases, and storage. - Strong analytical, troubleshooting, communication, and documentation skills with a collaborative mindset.
Nice to have - Experience supporting microservices-based architectures. - Depth across multiple observability platforms. - Experience with incident management, root cause analysis, and post-incident reviews. - Infrastructure as Code experience with Terraform, Ansible, or similar tools. - Familiarity with eBPF or low-level Linux performance monitoring. - Background building custom telemetry, ETL, or operational data pipelines. - Understanding of security monitoring, audit logging, and compliance requirements.
Benefits and work setup - Rotating on-call schedule with occasional after-hours or incident response responsibilities. - Hybrid or remote work arrangements depending on business needs. - Base salary range of $145,000 to $180,000 USD, with potential eligibility for bonus, equity, or commission. - Benefits package may include medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.