arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向动作的动作注意力:策略学习的涌现视觉瓶颈

Attention from Action, for Action: Emergent Visual Bottlenecks for Policy Learning

Zheyu Zhuang, Ruiyu Wang, Nick Heppert, Johannes Fabian Hahn, Abhinav Valada, Florian T. Pokorny, Danica Kragic

arXiv 2608.13422首次发表:更新:

发表机构

University of Freiburg; Universität Hamburg; KTH Royal Institute of Technology(弗赖堡大学; 汉堡大学; 瑞典皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人员提出 Seeker,一种从动作中学习注意力的任务状态条件化读出机制,可生成感知进展的 ROI,在仿真与真实环境中提升了视觉运动学习的数据效率和鲁棒性,真实机器人实验中成功率显著提高。

AI 中文摘要

将策略输入聚焦于感兴趣区域(ROI)的视觉瓶颈,可通过将观察位置与动作执行分离,提升数据高效的视觉运动学习。许多ROI接口依赖外部空间标签,如注视、物体类别或 affordance 标注;无标签替代方案通常通过检测夹爪或运动事件从轨迹中提取裁剪区域,并以投影末端执行器为中心固定裁剪。这类源自动作的裁剪是无需额外标签的有用空间先验,但编码了关于事件时序、代理点和裁剪尺度的固定选择。当控制所需视觉证据远离末端执行器或随任务进展连续变化时,这些裁剪可能错位。我们提出 Seeker,一种任务和状态条件化的读出机制,可从动作中学习注意力。从冻结的 DINOv3 特征开始,Seeker 用收集的视觉证据迭代更新查询,仅通过动作监督生成感知进展的 ROI。学习到的 ROI 用作 RGB 裁剪、掩码引导背景增强和点云过滤的空间接口。在仿真和真实环境中,Seeker 相比无裁剪、增强和动作衍生裁剪基线,提升了数据效率和鲁棒性;在真实机器人上,Seeker 将平均域内成功率从最佳基线的 48.3% 提升至 76.7%,在光照/背景变化下的成功率从 20.0% 提升至 60.0%。

英文摘要

Visual bottlenecks that focus policy inputs on regions of interest (ROIs) can improve data-efficient visuomotor learning by separating where to look from how to act. Many ROI interfaces rely on external spatial labels, such as gaze, object classes, or affordance annotations. Label-free alternatives often derive crops from trajectories by detecting gripper or motion events and centering a fixed crop at the projected end-effector. Such action-derived crops are useful spatial priors that require no additional labels, but they encode fixed choices about event timing, proxy points, and crop scale. When the visual evidence needed for control lies away from the end-effector or changes continuously with task progress, these crops can become misaligned. We propose Seeker, a task- and state-conditioned readout that learns attention from action. Starting from frozen DINOv3 features, Seeker iteratively updates a query with gathered visual evidence, producing progression-aware ROIs solely from action supervision. The learned ROI serves as a spatial interface for RGB cropping, mask-guided background augmentation, and point-cloud filtering. In simulation and the real world, Seeker improves data efficiency and robustness over no-crop, augmentation, and action-derived crop baselines. On real robots, Seeker raises average in-domain success from the best baseline's 48.3% to 76.7% and success under lighting/background shifts from 20.0% to 60.0%.

Comments[CoRL 2026] Code: https://github.com/zheyu-zhuang/seeker

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑