arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ActiveWAM:面向世界-动作模型的证据感知主动视觉

ActiveWAM: Evidence-Aware Active Vision for World-Action Models

Renjun Wu, Luzhou Ge, Xuesong Li

arXiv 2610.01698首次发表:更新:

发表机构

Beijing Institute of Technology(北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ActiveWAM提出证据感知的保持-获取问题,通过训练时反转统一学习观测与操作,在多个基准上显著提升成功率。

AI 中文摘要

主动视觉操作要求策略同时控制其相机和末端执行器,然而相机运动决定了在有限观测窗口内哪些证据保持可见。获取新视角可能会使任务关键线索移出视野,而保持当前视角则会错过潜在有用的观测。我们将此问题形式化为一个证据感知的保持-获取问题,并提出了ActiveWAM,一个统一的世界-动作模型,该模型联合学习观测与操作。为此,我们提出了训练时反转方法,该方法通过任务相关的源证据和可见的时间变化来约束一个冻结的视频先验,从而消除了测试时反转或候选排序的需要。在部署时,策略从视角感知的历史中生成双臂和云台动作(包括保持和重新获取行为),并根据新测量的RGB观测更新上下文。未来视频预测作为协同训练信号,而动作生成既不需要未来视频解码,也不需要最优视角标注。我们引入了RoboTwin-AV,一个包含50个任务的基准,具有可执行的云台控制与自动生成的演示。ActiveWAM在TAVIS分布外成功率上比最强基线提高了最多17.0个百分点,在RoboTwin-AV上比Fast-WAM提高了20.0个百分点,并在真实世界物理厨房任务上比Fast-WAM高出26.7个百分点。

英文摘要

Active vision manipulation requires a policy to control both its camera and its end-effectors, yet camera motion determines which evidence remains visible within finite observation windows. Acquiring a new view can displace task-critical cues, while retaining a view forgoes potentially useful observations. We formulate this as an evidence-aware retain--acquire problem and present ActiveWAM, a unified world--action model that learns observation and manipulation jointly. To this end, we propose training-time inversion which constrains a frozen video prior by task-bearing source evidence and visible temporal changes, eliminating the need for test-time inversion or candidate ranking. At deployment, the policy generates bimanual and pan/tilt actions-including stay and reacquisition behaviors-from view-aware history, and updates context from newly measured RGB observations. Future-video prediction serves as a co-training signal, while action generation requires neither future-video decoding nor optimal viewpoint annotations. We introduce RoboTwin-AV, a 50-task benchmark with executable pan/tilt control and automatically generated demonstrations. ActiveWAM improves TAVIS out-of-distribution success by up to 17.0 percentage points over the strongest baselines, achieves 20.0 additional points over Fast-WAM on RoboTwin-AV, and outperforms it by 26.7 points on real-world physical kitchen tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑