Remote job
Software Engineer (Platform Engineering)
Job details
About this role
Role overview
A platform engineering role focused on the systems that observe large-scale GPU infrastructure. The work centers on keeping a monitoring platform fast, observable, and continuously available, with an emphasis on reliability and performance at scale.
Responsibilities
- Maintain and improve the performance of an infrastructure-monitoring platform built around GPU deployments - Strengthen observability across the platform, including metrics, logs, and tracing where applicable - Ensure high availability and rapid incident response for always-on monitoring services - Evolve the platform's architecture to support growing scale and new monitoring needs - Collaborate with stakeholders who depend on the platform for infrastructure visibility
Requirements
- Background in software engineering with platform, infrastructure, or SRE experience - Familiarity with observability tooling and practices for distributed systems - Comfort working on performance, reliability, and operational quality at scale - Ability to navigate trade-offs between speed, cost, and uptime in platform design
Nice to have
- Hands-on experience with GPU or compute infrastructure environments - Exposure to monitoring pipelines that handle high-volume telemetry
Source material for this posting was limited, so the description focuses on the platform-engineering mission stated in the listing rather than introducing additional details.