← Back to jobs

Remote job

Data Engineer

wynd-labs

Data Engineer Full-time Permanent

Job details

Not specified Salary
Remote Eligibility
Not specified Experience
Full-time Employment

About this role

Role overview A data engineer is needed to help operate and extend the web-scale infrastructure that delivers massive volumes of public web data to teams training frontier AI models. The role centers on maintaining distributed crawlers, ingestion and annotation pipelines, and analytical data systems that power dataset creation for machine learning research. It is a fully remote position on a lean, technically rigorous team that values ownership, low ego, and fast iteration.

Responsibilities - Maintain, optimize, and troubleshoot database queries and supporting data systems to ensure reliable, efficient access and processing. - Build, maintain, and improve data pipelines that collect, process, transform, validate, and deliver large-scale datasets. - Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts and tools that gather publicly available data. - Monitor pipelines for failures and data quality issues, then drive timely fixes to preserve accuracy and operational continuity. - Document engineering work including queries, pipeline processes, scraping workflows, technical decisions, and resolutions. - Contribute to research and development projects that improve internal data products and workflows.

Requirements - Bachelor's degree or equivalent practical experience. - Advanced Python skills, including async programming, multiprocessing, and production-grade code for long-running data jobs. - Hands-on experience with high-volume web scraping, including proxies, rate limiting, anti-bot evasion, and platform APIs. - Experience designing and operating distributed data pipelines using task queues such as Celery, Kafka, or RabbitMQ. - Practical experience with columnar or analytical warehouses such as Databend, ClickHouse, or BigQuery, including partitioning and cost-aware querying. - Comfort with Docker and Kubernetes, including writing Helm charts, managing deployments, and autoscaling workloads. - Linux and bare-metal operations skills, including debugging disk I/O, network, and memory performance issues without managed-cloud abstractions. - Experience with CI/CD for data workflows using GitHub Actions or ArgoCD, plus building scalable APIs.

Benefits and work setup - Fully remote team with a lean, high-output engineering culture. - Competitive salary, benefits, and equity package.

Skills detected in the listing

PythonData WarehousingData EngineeringDockerKubernetes
Detected Sep 3, 2026
Last verified Sep 3, 2026
Original source: jobs.ashbyhq.com · application link requires Hidden Jobs Access

Hidden Jobs Access

Unlock application links

Read the full job details for free. An active Hidden Jobs Access subscription is required to open the original application link.

Weekly

FREE $6.99/week after trial
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
  • Cancel anytime before day 7

Monthly

$35.99 $17.99 /month
  • 35% cheaper than weekly
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching

Lifetime

$99.99 $49.99 /forever
  • One-time payment
  • Original application links
  • Instant job alerts
  • Premium filters and CV matching
Hidden Jobs gives subscribers direct access to original application links
Offer ends in 00:00:00 Your profile-fit rate expires at midnight