arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboRecover:在执行偏差下的机器人策略恢复基准

RoboRecover: Benchmarking Robot Policy Recovery under Execution Deviations

Yang Li, Chen Zhao, Zhuoran Wang, Jiankang Wang, Chao Shao, Yihan Lin, Haitao Shen, Jing Zhang

arXiv 2609.28952首次发表:更新:

发表机构

School of Information, Renmin University of China; Key Laboratory of Data Engineering and Knowledge Engineering; University of Science and Technology of China; Engineering Research Center of Database and Business Intelligence(中国人民大学信息学院; 数据工程与知识工程重点实验室; 中国科学技术大学; 数据库与商务智能工程研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RoboRecover基准通过重放动作前缀重建执行偏差状态,评估策略恢复能力,发现初始性能不决定恢复性能,确立恢复为独立评估维度。

AI 中文摘要

机器人策略基准日益覆盖多样化的任务和预设的分布外条件,但通常从预定义的初始状态评估完整轨迹。这些评估往往关注初始化的场景和最终结果,而对动态交互过程的关注较少。在闭环执行过程中,动作和接触可能改变物体关系和任务进度,产生需要恢复的非标称中间状态。恢复要求策略推断任务进度如何变化,纠正相关关系,并继续原始目标。我们引入RoboRecover,一个在执行偏差下进行机器人策略恢复的基准。RoboRecover从轨迹中选择偏差状态,通过重放动作前缀来重建它们,并在原始任务上评估策略。RoboRecover包含RoboTwin和LIBERO上的2,000个场景,每个平台有1,000个场景和固定的800/200训练/测试划分。结果表明,初始状态性能不能决定恢复性能,且策略在不同场景中表现出不同的恢复强度。利用其训练划分,RoboRecover进一步支持恢复干预的研究。RoboRecover将执行引起的中间状态恢复确立为机器人策略评估的一个独立维度。

英文摘要

Robot-policy benchmarks increasingly cover diverse tasks and preset out-of-distribution conditions, but typically evaluate complete trajectories from predefined initial states. These evaluations often focus on the initialized scene and the final outcome, while paying less attention to the dynamic interaction process. During closed-loop execution, actions and contacts can alter object relations and task progress, producing off-nominal intermediate states that need recovery. Recovery requires a policy to infer how task progress has changed, correct the relevant relations, and continue the original goal. We introduce RoboRecover, a benchmark for robot policy recovery under execution deviations. RoboRecover selects deviation states from trajectories, reconstructs them by replaying action prefixes, and evaluates policies on the original task. RoboRecover contains 2,000 scenarios across RoboTwin and LIBERO, with 1,000 scenarios and a fixed 800/200 train/test split on each platform. Results show that initial-state performance does not determine recovery performance and policies exhibit different recovery strengths across scenarios. Using its training split, RoboRecover further supports study on recovery interventions. RoboRecover establishes recovery from execution-induced intermediate states as a distinct dimension of robot policy evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑