Remote job
Senior Software Engineer - Observability / Runtime Systems
Job details
About this role
Role overview A senior engineering role centered on building the observability backend that supports large-scale GPU and AI runtime telemetry. The work focuses on high-performance systems responsible for ingesting, processing, and making sense of massive telemetry streams. It is a deep systems position at the intersection of distributed infrastructure, performance engineering, and AI runtime data.
Responsibilities - Design and build high-performance services that ingest GPU and AI runtime telemetry at scale - Develop backend systems that transform raw telemetry into structured, queryable data - Optimize ingestion and processing pipelines for throughput, latency, and reliability - Shape the runtime data plane so observability keeps pace with telemetry volume - Partner with infrastructure and AI teams to align backend design with runtime needs
Requirements - Senior-level experience in backend or systems software engineering - Proficiency in a high-performance systems language such as C++, Rust, or Go - Background designing distributed data pipelines or stream processing systems - Familiarity with GPU computing environments and AI/ML runtime concepts - Comfort working at the boundary between low-level performance work and platform engineering