arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以实例为锚的交互证据:将机器人规划建立在人类指向与操作的基础上

Instance-anchored interaction evidence: Grounding robot plans in human pointing and handling

Xinliang Xiao, Bowen Yang, Wenjing Zhang, Li Yang, Wei Zhou

arXiv 2610.12157首次发表:更新:

发表机构

Nanjing University of Science and Technology(南京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出以实例为锚的交互证据(IAE)方法,结合反向掩码传播、证据网络与语法约束动态程序,在WatchAct基准任务上显著提升机器人规划成功率,优于32B视觉语言模型等对照方法。

AI 中文摘要

协助人类的机器人常常需要依据人类展示的内容而非口头指令行动:例如多个相同纸箱中哪一个被指向,或哪个箱子被操作。机器人的规划从最终场景执行,而交互证据出现在更早的时刻,可能涉及已移动的物体。我们提出以实例为锚的交互证据(Instance-anchored Interaction Evidence, IAE),该方法将最终场景中的每个物体与其公开标识符关联,通过反向掩码传播在视频中保留每个物体的身份,并通过手、前臂与这些实例之间的几何关系描述每一帧。仅从任务结果训练的证据网络对实例进行评分。对于指向任务,采用结构化损失训练的语法约束动态程序解码对象-目标规划;符号程序处理引用消歧,且无需学习即可处理情景任务。在1255个WatchAct基准请求上,通过符号执行评分,IAE在隐含意图任务上的规划成功率达64.2%,而32B视觉语言模型的成功率为27.5%(严格成功率为49.7%,对比15.4%);在无特定任务训练的恢复、反转和模仿任务上,IAE的成功率为46.4%,对比27.0%。具有相同感知覆盖、相同32帧、正向跟踪或在相同标签上训练的关系模型的对照组无法解释该提升。将IAE的证据作为文本并解释其含义后,同一语言模型的成功率达57.4%:大部分提升来自以实例为锚的证据,而显式程序以极低成本额外提升6.8个百分点。指向任务仍是最困难的情况,严格成功率为16.9%。代码可在该https URL获取。

英文摘要

A robot that assists people must often act on what a person has shown rather than said: which of several identical cartons was pointed at, or which box was handled. The plan is executed from the final scene, whereas the evidence occurs earlier, possibly on objects that have since moved. We propose instance-anchored interaction evidence (IAE), which registers every object of the final scene to its public identifier, keeps each identity through the video by backward mask propagation, and describes every frame by the geometry between hands, forearms and these instances. An evidence network trained only from task outcomes scores the instances. For pointing tasks, a grammar-constrained dynamic program trained with a structured loss decodes object-destination programs; symbolic programs handle reference disambiguation and, without learning, episodic tasks. On 1,255 WatchAct benchmark requests, scored by symbolic execution, IAE reaches 64.2% plan success on implicit-intent tasks against 27.5% for a 32B vision-language model (strict success 49.7% against 15.4%), and 46.4% against 27.0% on restoration, reversal and imitation without task-specific training. Controls with the same perception overlays, the same 32 frames, forward tracking, or a relation model trained on the same labels do not explain the gain. Given IAE's evidence as text with its meaning explained, the same language model reaches 57.4%: most of the gain comes from the instance-anchored evidence, and the explicit programs add 6.8 points at a fraction of the cost. Pointing remains the hardest case, with 16.9% strict success. The code is available at https://github.com/WeiZhou96/iae-watchact.

Comments22 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑