Remote job
Software Engineer | Observability
Job details
About this role
Role overview
Build the observability platform powering a real-time, cloud-native data platform used by enterprises for high-performance transactional and analytical workloads. This hands-on engineering role owns features end-to-end at the intersection of distributed systems, cloud infrastructure, database technology, and AI-powered observability, helping customers understand usage patterns, identify bottlenecks, optimize performance, and manage alerting across traces, logs, and metrics. The position partners closely with Product and customer-facing teams to deliver meaningful business impact in a continuously shipping environment.
Responsibilities
- Design and implement scalable observability features for traces, logs, and metrics spanning ingestion, processing, storage, and visualization. - Work across control plane and data plane components in a multi-cloud environment (AWS, GCP, Azure), ensuring reliable operation and data consistency at scale. - Build high-throughput data pipelines that process telemetry data using OpenTelemetry Collector and related open-source tooling. - Develop and maintain alerting capabilities with Alertmanager, enabling customers to define, tune, and manage alerts with routing, inhibition, and notification management. - Optimize time-series data storage and query performance, handling high-cardinality data and complex analytical queries. - Contribute to data visualization dashboards in Grafana, creating intuitive customer-facing experiences for exploring telemetry data. - Collaborate closely with Product Management to translate customer and business requirements into robust technical solutions. - Investigate and resolve difficult production issues, debugging data synchronization across distributed systems and cloud providers. - Participate in on-call rotations to ensure system reliability and respond to incidents promptly.
Requirements
- 2+ years of professional software development experience building distributed systems or backend services. - Strong proficiency in Go (Golang); experience with Rust, Python, or C++ is also valuable. - Deep understanding of distributed systems concepts including scalability, consistency, high availability, concurrency, and failure modes. - Familiarity with distributed systems managed via Kubernetes. - Demonstrated ability to design and build reliable, high-performance system software. - Familiarity with observability concepts such as traces, logs, metrics, APM, and monitoring patterns. - Strong problem-solving and debugging skills with the ability to root-cause complex production issues. - Excellent written and verbal communication skills, with ability to collaborate in multicultural, remote-first teams.
Nice to have
- Experience with time-series data and understanding of metrics cardinality challenges. - Proficiency with SQL and experience working with relational or distributed databases. - Experience building cloud-native SaaS platforms with multi-tenant architecture. - Hands-on experience with Grafana, Alertmanager, Loki, Tempo, OpenTelemetry Collector, and OTLP protocol. - OpenTelemetry expertise or active contributions to OTel projects. - Experience with time-series databases such as Prometheus TSDB, InfluxDB, Mimir, TimescaleDB, or distributed SQL databases. - Experience with data pipeline technologies including Apache Kafka, Parquet, Arrow, or stream processing frameworks. - Experience working with AI agents or LLM-powered applications, including agentic workflows for observability.