Job details
About this role
Role overview A data engineer is needed to help operate and extend the web-scale infrastructure that delivers massive volumes of public web data to teams training frontier AI models. The role centers on maintaining distributed crawlers, ingestion and annotation pipelines, and analytical data systems that power dataset creation for machine learning research. It is a fully remote position on a lean, technically rigorous team that values ownership, low ego, and fast iteration.
Responsibilities - Maintain, optimize, and troubleshoot database queries and supporting data systems to ensure reliable, efficient access and processing. - Build, maintain, and improve data pipelines that collect, process, transform, validate, and deliver large-scale datasets. - Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts and tools that gather publicly available data. - Monitor pipelines for failures and data quality issues, then drive timely fixes to preserve accuracy and operational continuity. - Document engineering work including queries, pipeline processes, scraping workflows, technical decisions, and resolutions. - Contribute to research and development projects that improve internal data products and workflows.
Requirements - Bachelor's degree or equivalent practical experience. - Advanced Python skills, including async programming, multiprocessing, and production-grade code for long-running data jobs. - Hands-on experience with high-volume web scraping, including proxies, rate limiting, anti-bot evasion, and platform APIs. - Experience designing and operating distributed data pipelines using task queues such as Celery, Kafka, or RabbitMQ. - Practical experience with columnar or analytical warehouses such as Databend, ClickHouse, or BigQuery, including partitioning and cost-aware querying. - Comfort with Docker and Kubernetes, including writing Helm charts, managing deployments, and autoscaling workloads. - Linux and bare-metal operations skills, including debugging disk I/O, network, and memory performance issues without managed-cloud abstractions. - Experience with CI/CD for data workflows using GitHub Actions or ArgoCD, plus building scalable APIs.
Benefits and work setup - Fully remote team with a lean, high-output engineering culture. - Competitive salary, benefits, and equity package.