Remote job
Senior Data Engineer
Job details
About this role
Role overview
A senior engineer is needed to take full ownership of a global edge data pipeline that captures real-world events, normalizes and timestamps them at ingestion, and delivers evaluated records to downstream consumers such as trading desks and ML models. The work sits at the intersection of streaming infrastructure and data quality, where freshness, latency, and provenance directly determine what customers can trust. It is high-scale, production-grade systems work on a small team with broad technical scope.
Responsibilities
- Design, deploy, and operate a worldwide listening layer that ingests continuously from many external sources and timestamps each observation at the receiving node. - Build both streaming and bulk-delivery paths so records reach customers with minimal added latency. - Maintain point-in-time history, deduplication, and end-to-end provenance so downstream users can rely on what they see. - Engineer reliability at scale through backpressure handling, retries, and clean recovery when upstream sources misbehave. - Partner with the ML team to ensure the ingestion layer hands models clean, well-formed input on time. - Monitor production pipelines, catch freshness or quality regressions early, and protect customer-facing reliability.
Requirements
- At least five years building and operating data-intensive systems in production, with real ownership of pipelines that served live traffic. - Strong Python and SQL, plus hands-on fluency with a streaming stack such as Kafka or Flink and modern data tooling like Spark, Arrow, or Airflow. - Experience with high-volume ingestion and web-scale collection, including the reliability problems those workloads create. - Comfort with the production stack: containers, Kubernetes, CI/CD, and a major cloud such as AWS or GCP. - A practical feel for latency and slow-tail behavior, not just average throughput. - Comfort owning open-ended production problems on a small team with broad technical scope.
Nice to have
- Background supporting ML or model-training pipelines with carefully governed inputs. - Prior work on global, geographically distributed ingestion systems where freshness and provenance are first-class concerns.