Remote job
Research Engineer, Benchmarking, Robotics
Job details
About this role
Role overview
This role designs the benchmarks and evaluation systems used to measure the capabilities of general-purpose robot models. The work goes beyond running predefined tests: it involves deciding which abilities matter, how experiments should be controlled, which metrics are meaningful, and how to distinguish model performance from effects caused by the setup. Evaluations will span simulation and physical robots and should produce reproducible evidence that guides research decisions.
Responsibilities
- Design benchmark suites for robot foundation models and policies covering manipulation, perception, reasoning, adaptation, recovery, and long-horizon behavior. - Define evaluation protocols, success criteria, quantitative metrics, controlled variations, and trial requirements for simulation and real-robot experiments. - Test generalization across unseen objects, environments, tasks, embodiments, instructions, poses, scenes, and initial conditions. - Build automated or semi-automated evaluation infrastructure for comparing models and checkpoints, including tools for recording trajectories, video, observations, actions, failures, and metadata. - Create stress tests, failure taxonomies, diagnostic tools, dashboards, and visualizations that reveal regressions, edge cases, uncertainty, and performance variance. - Establish evaluation gates and communicate findings so that benchmark results influence model architecture, training data, and research priorities.
Requirements
- Strong software engineering and research skills, with the ability to turn questions about robot capability into controlled, reproducible experiments. - Experience designing quantitative metrics, experimental protocols, or evaluation systems where measurements may be noisy or affected by initial conditions. - Understanding of robotics and the factors that influence physical performance, including embodiment, environment, perception, control, latency, and object variation. - Ability to analyze failures systematically and separate meaningful model changes from artifacts of the evaluation procedure. - Comfortable collaborating with robotics engineers and researchers across simulation, hardware, machine learning, and data systems.
Nice to have
- Experience evaluating robot-learning or Vision-Language-Action models, designing manipulation benchmarks, or building benchmark datasets and automated evaluation pipelines. - Familiarity with ROS or ROS2; simulation platforms such as MuJoCo, Isaac Sim, ManiSkill, or RLBench; or hardware-in-the-loop testing. - Experience with large-scale robot experiments, statistical analysis of physical trials, and publications or open-source work in robotics evaluation, benchmarks, robot learning, or embodied AI.
What success looks like
Researchers can compare model versions without relying on selected demonstrations, break improvements down by capability, quantify and categorize failures, and reproduce real-world results reliably enough to support decisions. Benchmarks are challenging enough to expose meaningful differences and closely connected to useful physical robot behavior.