Remote job
Senior Software Engineer - Observability / Runtime Systems
Job details
About this role
Role overview
This senior engineering role focuses on the telemetry and observability systems that make GPU and AI workloads legible to operators and security teams. The work centers on high-throughput pipelines that collect, process, and surface signals from accelerated-compute environments at scale. It is a backend-heavy position with a strong emphasis on systems performance and reliability.
Responsibilities
- Design and build the observability backend that ingests telemetry from GPU and AI runtime environments - Develop high-performance processing pipelines for large-scale time-series and event data - Implement query, storage, and aggregation layers that surface operational and security insights - Collaborate with research and platform teams to define instrumentation standards across the stack - Optimize data paths for throughput, latency, and cost as data volumes grow - Maintain reliability of the platform as query complexity and ingestion rates increase
Requirements
- Senior-level experience building high-performance backend, data, or distributed systems - Strong proficiency in systems languages such as Rust, Go, or C++ - Familiarity with observability tooling, telemetry pipelines, and time-series storage systems - Understanding of GPU compute, AI runtimes, or other accelerated-compute environments - Comfort designing for scale, including streaming, sharding, and backpressure handling