← Back to jobs

Remote job

Principal Observability Platform Engineer

DevOps Full-time Permanent United States

Job details

$190,000—$300,000 Salary
United States Eligibility
Principal Experience
Full-time Employment

About this role

Role overview

Set the technical direction and build the observability platform used to understand large-scale infrastructure, accelerated computing clusters, and AI workloads. This principal-level role treats observability as a core engineering product: you’ll shape its architecture, improve how teams detect and diagnose failures, and guide other engineers toward systems that remain clear and manageable as they grow.

Responsibilities - Define platform architecture and long-term strategy for metrics, logs, traces, and alerting. - Make durable decisions about tools, data models, ingestion, retention, and management of high-cardinality data. - Find gaps in visibility before they lead to incidents and design systems that make problems easier to identify and diagnose. - Partner with site reliability, infrastructure, and AI/ML teams to build observability into their systems and working practices. - Establish engineering patterns that teams can adopt, mentor observability engineers, and review designs and implementation. - Lead incident reviews and turn findings into lasting platform improvements. - Assess tools for signal quality, scale, and operational efficiency, introducing useful capabilities and retiring unsuitable ones.

Requirements - Eight or more years in SRE, infrastructure, platform engineering, or observability-focused work. - Experience operating observability infrastructure at substantial scale and designing for significant growth. - Deep hands-on knowledge of several tools in the observability ecosystem, such as Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, or Elastic. - Strong engineering skills in Python, Go, or a similar language, with the ability to own complex systems end to end. - Kubernetes experience at scale and infrastructure-as-code practice using Terraform, Ansible, or an equivalent tool. - Ability to communicate design trade-offs clearly and influence teams through technical judgment and collaboration.

Nice to have

Experience with GPU or HPC environments such as Slurm; high-volume data pipelines using tools such as Kafka, Vector, or Fluent Bit; AI/ML observability covering GPU use, training jobs, or inference latency; or organization-wide observability strategy.

Benefits and work setup

The listed base salary range is $190,000–$300,000 USD. The role may also include bonus, equity, or commission eligibility. Potential benefits include medical, dental, and vision coverage, flexible paid time off, parental leave, and retirement plan participation.

Skills detected in the listing

PythonGoKubernetesTerraform
Detected Oct 9, 2026
Last verified Oct 9, 2026

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Daily or weekly job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight