arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关注关键区域:面向视觉-语言-动作模型的自适应视觉细化

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models

Jin Cui, Yanbin Hu, Xinyue Long, Linkai Li, Boran Zhao, Pengju Ren

arXiv 2608.02197首次发表:更新:

AI 中文总结

针对VLA模型视觉表示不可靠的问题,提出AtVLA框架,结合注意力校正与不确定性门控局部细化,在多基准测试中显著提升机器人操作成功率且计算开销可控。

AI 中文摘要

视觉-语言-动作(VLA)模型的视觉表示对于空间精确的机器人操作仍不可靠。我们发现,VLA中的视觉编码器也存在通用视觉Transformer中已记录的注意力伪影,且进一步表明,在具身策略中,这些伪影与后训练期间获得的空间感知能力密切相关。当编码器学习任务相关信息(如物体位置、深度顺序和局部几何)时,有限的全局token容量会导致部分信息溢出到低信息patch token中。我们引入AtVLA,一种向视觉编码器插入可学习register token的框架。仅使用具身数据和原始动作目标进行端到端训练后,这些register成为具身空间信息的专用载体,而其余patch token则恢复干净且空间保真的注意力分布,这对精确的目标定位和细粒度接触至关重要。干净的注意力恢复了可靠的定位,但无法恢复低分辨率观测中丢失的几何细节。因此,AtVLA将注意力校正与不确定性门控局部细化相结合:动作专家采样多个动作chunk并从其分歧中估计不确定性;仅对不确定的预测,动作条件注意力rollout识别任务相关区域,该区域被裁剪、以高分辨率重新编码,并附加到缓存前缀中以生成细化动作。在LIBERO、SimplerEnv和具有挑战性的单视图真实世界基准测试中,AtVLA将LIBERO平均成功率从94.2%提升至98.4%,真实世界成功率从46.5%提升至69.0%。裁剪在约30%的重规划步骤中触发,在代表性部署设置下,总计算量仅为基础模型的1.4-1.6倍。

英文摘要

Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover that vision encoders in VLAs also exhibit attention artifacts previously documented in generic Vision Transformers, and further show that, in embodied policies, these artifacts are closely associated with spatial perception capabilities acquired during post-training. As the encoder learns task-relevant information such as object location, depth ordering, and local geometry, limited global-token capacity causes part of this information to spill into low-information patch tokens. We introduce AtVLA, a framework that inserts learnable register tokens into the visual encoder. Trained end-to-end using only embodied data and the original action objective, these registers emerge as dedicated carriers of embodied spatial information, while the remaining patch tokens recover clean and spatially faithful attention distributions crucial for precise target localization and fine-grained contact. Clean attention restores reliable localization, but cannot recover geometric details lost in low-resolution observations. AtVLA therefore couples attention rectification with uncertainty-gated local refinement. The action expert samples multiple action chunks and estimates uncertainty from their disagreement; only for uncertain predictions, action-conditioned attention rollout identifies the task-relevant region, which is cropped, re-encoded at high resolution, and appended to the cached prefix for refined action generation. Across LIBERO, SimplerEnv, and a challenging single-view real-world benchmark, AtVLA improves the average LIBERO success rate from 94.2% to 98.4% and real-world success from 46.5% to 69.0%. The cropping is triggered on approximately 30% of replanning steps, resulting in only 1.4-1.6x the total computation of the base model under the representative deployment setting.

Comments13 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑