Remote job
Data Engineer, Embodied AI
Job details
About this role
Role overview
This role builds the data infrastructure that supports embodied-AI policy learning. The focus is on generating, ingesting, processing, curating, storing, and serving multimodal experience from physical robots, simulation, demonstrations, existing datasets, and synthetic-generation systems. The work is oriented toward model training: improving data quality, coverage, reproducibility, and the speed at which new experience or policy failures become usable training data.
Responsibilities
- Build end-to-end data pipelines for embodied-AI training, from collection and ingestion through processing, storage, curation, and serving. - Design schemas and representations for multimodal trajectories, and synchronize video, images, actions, robot state, sensor data, language, metadata, and evaluation results. - Develop systems for filtering, cleaning, deduplicating, balancing, and curating datasets across tasks, robots, environments, and sources. - Create automated methods for identifying useful, difficult, failed, or unusual episodes and converting policy failures into candidate training data. - Establish dataset versioning, lineage, reproducibility, quality control, observability, and efficient loading for large-scale model training. - Build inspection and debugging tools, sampling and mixture strategies, and quantitative measures of dataset quality and coverage in partnership with ML, robotics, and simulation engineers.
Requirements
- Strong software and data-engineering fundamentals, including the ability to design reliable systems that evolve as training requirements change. - Experience building data pipelines and schemas for complex, multimodal, sequential, or trajectory-based data. - Strong understanding of data quality, reproducibility, lineage, and observability. - Ability to reason about data from a machine-learning perspective and connect dataset composition with policy performance. - Comfortable working closely with researchers and engineers to diagnose data problems and improve training workflows.
Nice to have
- Experience preparing datasets for deep learning or foundation-model training, including video, image, multimodal, sequential, or robot-trajectory data. - Background in embodied AI, robotics, autonomous systems, reinforcement learning, imitation learning, or learning from demonstrations. - Experience with large-scale dataset generation, synthetic-data pipelines, simulation-generated data, distributed training input pipelines, data labeling, active learning, or automated data selection.
What success looks like
Training engineers can quickly determine what data exists, where it came from, which tasks or conditions are underrepresented, which failures should be added, and whether a changed data mixture improved a policy. Simulation runs, physical experiments, evaluations, and policy failures form a practical data flywheel for improving subsequent models.