arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

事件对齐的视觉动作推理用于世界动作模型

Event-Aligned Visual Action Reasoning for World Action Models

Xiaomeng Yang, Yushu Wu, Yi Gao, Yuhao Lei, Xuan Zhang, Pu Zhao, Yanzhi Wang

arXiv 2610.09427首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出事件对齐的视觉动作推理框架,通过围绕交互事件组织视觉预测并引入执行有效性头,显著提升世界动作模型的动作生成性能,在DOMINO和RoboTwin基准上验证了有效性。

AI 中文摘要

世界动作模型(WAMs)利用未来视觉预测作为中间推理过程来引导动作生成。然而,现有的WAMs通常根据预定义的时间间隔来构建视觉想象,没有明确考虑任务关键交互与连接转换的不同作用。我们认为,有效的视觉前瞻应直接与任务相关的交互及其相应的推理需求对齐。为此,我们引入了一个事件对齐的视觉动作推理框架,该框架围绕交互事件组织视觉动作预测。通过事件对齐的视觉动作监督,WAM学会在每个想象的展开中生成事件对齐的视觉上下文,更加重视为动作生成提供信息的关键状态变化。这根据底层交互动态塑造了视觉推理的粒度,在任务关键事件周围进行详细推理,并通过连接转换进行更粗略的推进。此外,我们引入了一个执行有效性头,用于识别每个预测动作序列的有效部分,避免在分块推理期间产生冗余动作。实验表明,在DOMINO成功率上比基线提高了10.26个百分点,并在RoboTwin 2.0上取得了有竞争力的性能。它还能从DOMINO级别1迁移到级别2和3,而无需目标级别的适配。

英文摘要

World-Action Models (WAMs) utilize future visual prediction as an intermediate reasoning process to guide action generation. However, existing WAMs typically structure visual imagination according to predefined temporal intervals, without explicitly accounting for the different roles of task-critical interactions and connecting transitions. We argue that effective visual foresight should align directly with task-relevant interactions and their corresponding reasoning demands. To this end, we introduce an event-aligned visual action reasoning framework that organizes visual-action prediction around interaction events. Through event-aligned visual-action supervision, WAM learns to generate event-aligned visual context in each imagined rollout, placing greater emphasis on critical state changes that inform action generation. This shapes the visual reasoning granularity according to the underlying interaction dynamics, with detailed reasoning around task-critical events and coarser progression through connecting transitions. Furthermore, we introduce an execution validity head that identifies the valid portion of each predicted action sequence, avoiding redundant actions during chunked inference. Experiments demonstrate a 10.26 percentage point improvement in DOMINO success rate over baseline and competitive performance on RoboTwin 2.0. It also transfers from DOMINO Level 1 to Levels 2 and 3 without target-level adaptation.

CommentsProject Page: https://xiaomeng-yang.github.io/Event-aligned-WAM/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑