Remote job
Principal Observability Platform Engineer
Job details
About this role
Role overview
Set the technical direction and build the observability platform used to understand large-scale infrastructure, accelerated computing clusters, and AI workloads. This principal-level role treats observability as a core engineering product: you’ll shape its architecture, improve how teams detect and diagnose failures, and guide other engineers toward systems that remain clear and manageable as they grow.
Responsibilities - Define platform architecture and long-term strategy for metrics, logs, traces, and alerting. - Make durable decisions about tools, data models, ingestion, retention, and management of high-cardinality data. - Find gaps in visibility before they lead to incidents and design systems that make problems easier to identify and diagnose. - Partner with site reliability, infrastructure, and AI/ML teams to build observability into their systems and working practices. - Establish engineering patterns that teams can adopt, mentor observability engineers, and review designs and implementation. - Lead incident reviews and turn findings into lasting platform improvements. - Assess tools for signal quality, scale, and operational efficiency, introducing useful capabilities and retiring unsuitable ones.
Requirements - Eight or more years in SRE, infrastructure, platform engineering, or observability-focused work. - Experience operating observability infrastructure at substantial scale and designing for significant growth. - Deep hands-on knowledge of several tools in the observability ecosystem, such as Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, or Elastic. - Strong engineering skills in Python, Go, or a similar language, with the ability to own complex systems end to end. - Kubernetes experience at scale and infrastructure-as-code practice using Terraform, Ansible, or an equivalent tool. - Ability to communicate design trade-offs clearly and influence teams through technical judgment and collaboration.
Nice to have
Experience with GPU or HPC environments such as Slurm; high-volume data pipelines using tools such as Kafka, Vector, or Fluent Bit; AI/ML observability covering GPU use, training jobs, or inference latency; or organization-wide observability strategy.
Benefits and work setup
The listed base salary range is $190,000–$300,000 USD. The role may also include bonus, equity, or commission eligibility. Potential benefits include medical, dental, and vision coverage, flexible paid time off, parental leave, and retirement plan participation.