Remote job
Principal Observability Platform Engineer
Job details
About this role
Role overview This is a senior, hands-on leadership role owning the technical direction of an observability platform that provides deep visibility into GPU clusters, AI workloads, and the underlying infrastructure. Observability is treated as a product and discipline, with the role setting architectural direction, defining standards, and ensuring the platform scales ahead of the business. The position combines deep engineering, platform design, technical mentorship, and incident leadership across metrics, logs, traces, and alerting.
Responsibilities - Own the technical strategy and architecture for observability across metrics, logs, traces, and alerting at significant scale. - Drive multi-year platform decisions covering tooling selection, data models, ingestion patterns, retention, and cardinality management. - Identify systemic reliability gaps before they become incidents and design platforms that make failures visible and fast to diagnose. - Partner with SRE, infrastructure, and AI/ML teams to embed observability natively into how services are built and operated. - Define standards and patterns that engineers adopt because they are demonstrably better, not because of mandate. - Mentor the observability team, lead incident postmortems, and evaluate new tooling while retiring what no longer serves the platform.
Requirements - 8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles, including operating observability infrastructure at serious scale. - Deep hands-on experience with a significant subset of Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, and Elastic. - Strong engineering fundamentals, with proficiency in Python, Go, or a comparable language, and comfort owning complex systems end to end. - Experience with Kubernetes at scale; familiarity with GPU infrastructure or HPC environments such as Slurm is a strong plus. - Infrastructure-as-Code as a default practice, using Terraform, Ansible, or equivalent tools.
Nice to have - Experience with high-volume streaming pipelines such as Kafka, Vector, or Fluent Bit. - Background in AI/ML infrastructure observability, including GPU utilisation, training-job visibility, and inference latency. - Prior experience defining organisation-wide observability strategy.
Benefits and work setup - Base salary range of $190,000-$300,000 USD, with potential eligibility for bonus, equity, or commission programs. - Benefits package may include medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.