arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00908cs.RO

何时停止:基于视觉-语言动作(VLA)内部交叉注意力动力学的自适应动作分块

Knowing When to Stop: Adaptive Action Chunking via Internal Cross-Attention Dynamics in VLAs

Runze Xu, Xiaolong Shan, Shuang Dai, Yu Wang, Jincheng Yu

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对 VLA 框架中固定动作分块时域的效率-准确性权衡问题,提出基于动作专家内部交叉注意力动力学的自适应分块方法,在多机器人操作任务上提升了任务成功率并保持了控制效率。

中文摘要 AI 辅助

动作分块是现代视觉-语言动作(Vision-Language-Action,VLA)框架中的标准执行策略,但固定执行时域会在效率与准确性之间产生权衡:短分块需要频繁推理,可能引发振荡行为;而长分块则可能与新观测到的状态失配。本文针对这一局限,提出了一种基于动作专家内部交叉注意力动力学的自适应动作分块方法。研究发现,随着预测时域延长,动作到观测的交叉注意力会愈发分散,其熵值会上升至一个平台期;该模式与更高的动作预测误差相关,可作为在线信号,指示当前观测为进一步开环执行提供的依据有限。基于此,本文引入一种无需训练的截断机制,用于检测持续的高熵平台期,并在推理阶段动态选择执行时域。该方法利用策略已计算的注意力权重,引入的额外开销可忽略不计。在 RoboTwin 2.0、LIBERO 及三个真实世界操作任务上,针对 π_{0.5} 和 X-VLA 的评估显示,与固定时域及自适应分块基线相比,该方法的平均任务成功率有所提升,同时保持了高效的闭环控制。这些结果表明,交叉注意力动力学可为 VLA 中的自适应动作执行提供实用的内部信号。

英文摘要

Action chunking is a standard execution strategy in modern Vision-Language-Action (VLA) frameworks, but fixed execution horizons impose a trade-off between efficiency and accuracy. Short chunks require frequent inference and may cause oscillatory behavior, whereas long chunks can become misaligned with newly observed states. We address this limitation with an adaptive action chunking approach based on internal cross-attention dynamics in the action expert. We observe that, as the prediction horizon extends, action-to-observation cross-attention becomes increasingly dispersed and its entropy rises toward a plateau. This pattern is associated with higher action prediction error and provides an online signal that the current observation offers limited grounding for further open-loop execution. Based on this observation, we introduce a training-free truncation mechanism that detects sustained high-entropy plateaus and dynamically selects the execution horizon during inference. The method uses attention weights already computed by the policy and introduces negligible additional overhead. Evaluations on $π_{0.5}$ and X-VLA across RoboTwin 2.0, LIBERO, and three real-world manipulation tasks show improved average task success over fixed-horizon and adaptive chunking baselines, while preserving efficient closed-loop control. These results show that cross-attention dynamics can provide a practical internal signal for adaptive action execution in VLAs.

发表机构

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑