发表机构
School of Augmented Intelligence; Arizona State University; NVIDIA(增强智能学院; 亚利桑那州立大学; 英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RoboReel基准,通过真实人类视频和模拟轨迹评估机器人观察学习技能,涵盖多算法测试,发现长时程和低容差任务仍是挑战。
AI 中文摘要
观察学习(LfO)是一种基础的机器人能力,它复制了人类和动物如何通过社交相互学习。除了其生物学上的相似性,这种模式为数据效率低下和数据匮乏的领域(如机器人学)中的数据扩展提供了实用解决方案。近期研究在从人类视频学习操作技能方面展示了有前景的结果,但该领域的进展仍然难以评估。现有方法在假设、硬件选择和环境设置上差异很大,使得难以进行有意义的比较并识别该领域的进展。为解决这些挑战,我们引入了RoboReel:一个用于评估从人类视频学习策略模型的统一基准。RoboReel包含捆绑的真实世界人类演示视频、模拟机器人轨迹以及十个操作任务的评估环境。我们开发了四个测试套件,以评估模型在多个维度上的性能,包括对视觉干扰物的鲁棒性和完成长时程任务的能力。我们的基准涵盖了不同类别的观察学习模型,并研究了多种表示选择在我们的基准评估中的有效性,该评估涵盖了LfO领域的七个以上最先进算法(包括我们基于VLA的变体)。最后,我们对不同类型的算法进行了分析,表明长时程任务和低容差任务对当前模型仍具挑战性。网页:此https URL
英文摘要
Learning from Observation (LfO) is a fundamental robotic capability that replicates how humans and animals socially learn from each other. Beyond its biological parallels, this modality provides a practical solution for data scaling in sample-inefficient and data-starved domains like robotics. Recent work has demonstrated promising results in learning manipulation skills from human videos, yet progress in this area remains difficult to assess. Existing methods vary widely in assumptions, hardware choices, and environment setups making it difficult to draw meaningful comparisons and identify advances in the field. To address these challenges, we introduce RoboReel: a unified benchmark for evaluating models that learn policies from human videos. RoboReel consists of bundled real-world human demonstration videos, simulated robot trajectories, and evaluation environments on ten manipulation tasks. We develop four test suites to evaluate the models' performance on multiple axes, including the robustness to visual distractors and the ability to complete long-horizon tasks. Our benchmark covers learning-from-observation models from different categories, and studies the effectiveness of multiple representation choices in our benchmark evaluation that covers over seven state-of-the-art algorithms (including our VLA based variants) in the field of LfO. Finally, we present an analysis of the different types of algorithms showing that long-horizon tasks and tasks with low tolerances are still challenging for current models. Webpage: https://roboreel.github.io
Comments31 pages, 8 tables, 11 figures. In Proceedings of CoRL 2026