从回滚到重置:用于自主长时程操作评估的基于图的方法
From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
- University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出HALTER,一种基于场景图和LLM的自主长时程操作评估重置方法,利用原子技能库规划重置,在Franka臂上显著提升重置成功率并减少操作员时间。
AI中文摘要:
机器人操作策略正在快速改进,而真实机器人评估仍然是证明这种进展的标准证据。然而,这种评估仍然依赖人类在每次回滚之间重置场景,这消耗了操作员的时间,并且使得初始状态分布不明确,导致结果难以复现。最近的一个系统AutoEval自动化了重置和评分,但仅适用于单步任务,因为长时程回滚可能终止于组合爆炸的众多配置,而单个学习到的重置策略无法覆盖所有这些配置。我们提出了HALTER,一个用于自主长时程任务评估和重置的方法,它通过规划一组学习到的原子重置技能来恢复场景,因此演示成本与技能库的大小成比例,而不是与终止状态的数量成比例。HALTER从点云和视觉基础模型在线构建空间场景图,并由大语言模型(LLM)基于该图进行推理,以对回滚进行评分、规划重置并验证重置是否成功,而无需为任何任务收集标记的成功图像。在Franka机械臂上的四个长时程任务中,HALTER在76%的情节中成功重置场景,而AutoEval为52%,运动规划重置为65%;HALTER在90%的情节中正确估计了已完成技能的比例,而AutoEval为76%。其重置验证结论在91%的情节中正确,而AutoEval为78%。此外,相对于手动重置,HALTER将评估活动的操作员时间减少了72%。我们进一步在三个保留任务上测量了组合泛化能力,其中HALTER重置了74.7%的情节,而针对每个任务的重置策略仅为1.3%,并且我们对场景表示和图更新率进行了消融研究。
英文摘要:
Robot manipulation policies are improving quickly, and real-robot evaluation remains the standard evidence for that progress. It still relies on a human to reset the scene between rollouts, which consumes operator time and leaves the initial state distribution unspecified, so results reproduce poorly. A recent system, AutoEval, automates both reset and scoring, but only for single-step tasks, because a long-horizon rollout can terminate in combinatorially many configurations that no single learned reset policy covers. We present HALTER, a Harness for Autonomous Long-horizon Task Evaluation and Reset, which restores the scene by planning over a library of learned atomic reset skills, so demonstration cost scales with the size of that library rather than with the number of terminal states. HALTER builds a spatial scene graph online from point clouds and vision foundation models, and an LLM reasons over this graph to score the rollout, plan the reset, and verify that the reset succeeded, without collecting labeled success images for any task. On four long-horizon tasks on a Franka arm, HALTER restores the scene in 76% of episodes, against 52% for AutoEval and 65% for a motion-planning reset, and it estimates the completed-skill fraction correctly in 90% of episodes, against 76%. Its reset-verification verdict is correct in 91% of episodes, compared with 78% for AutoEval. It also cuts the operator time of an evaluation campaign by 72% relative to manual reset. We further measure compositional generalization on three held-out tasks, where HALTER resets 74.7% of episodes against 1.3% for a per-task reset policy, and we ablate the scene representation and the graph update rate.