← Back to jobs

Remote job

Principal Observability Platform Engineer

DevOps Full-time Permanent US

Job details

$190,000—$300,000 Salary
US Eligibility
Principal Experience
Full-time Employment

About this role

Role overview This is a senior, hands-on leadership role owning the technical direction of an observability platform that provides deep visibility into GPU clusters, AI workloads, and the underlying infrastructure. Observability is treated as a product and discipline, with the role setting architectural direction, defining standards, and ensuring the platform scales ahead of the business. The position combines deep engineering, platform design, technical mentorship, and incident leadership across metrics, logs, traces, and alerting.

Responsibilities - Own the technical strategy and architecture for observability across metrics, logs, traces, and alerting at significant scale. - Drive multi-year platform decisions covering tooling selection, data models, ingestion patterns, retention, and cardinality management. - Identify systemic reliability gaps before they become incidents and design platforms that make failures visible and fast to diagnose. - Partner with SRE, infrastructure, and AI/ML teams to embed observability natively into how services are built and operated. - Define standards and patterns that engineers adopt because they are demonstrably better, not because of mandate. - Mentor the observability team, lead incident postmortems, and evaluate new tooling while retiring what no longer serves the platform.

Requirements - 8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles, including operating observability infrastructure at serious scale. - Deep hands-on experience with a significant subset of Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, and Elastic. - Strong engineering fundamentals, with proficiency in Python, Go, or a comparable language, and comfort owning complex systems end to end. - Experience with Kubernetes at scale; familiarity with GPU infrastructure or HPC environments such as Slurm is a strong plus. - Infrastructure-as-Code as a default practice, using Terraform, Ansible, or equivalent tools.

Nice to have - Experience with high-volume streaming pipelines such as Kafka, Vector, or Fluent Bit. - Background in AI/ML infrastructure observability, including GPU utilisation, training-job visibility, and inference latency. - Prior experience defining organisation-wide observability strategy.

Benefits and work setup - Base salary range of $190,000-$300,000 USD, with potential eligibility for bonus, equity, or commission programs. - Benefits package may include medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Skills detected in the listing

PythonGoKubernetesTerraform
Detected Oct 2, 2026
Last verified Oct 6, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight