arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09925cs.AIcs.RO

时间-频率几何交叉注意力用于分块视觉-语言-动作模型

Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

Shengye Dong, Haochen Niu, Hao Liu, Peiwen Lin, Chuang Wang, Shanmin Pang

首次发表
浏览论文内容

中文总结 AI 辅助

提出时间-频率几何交叉注意力(TFGCA)模块,通过小波分解和融合点积与楔积的交叉注意力,修复VLA模型在频率和跨相位几何上的盲点,提升分布内外成功率。

中文摘要 AI 辅助

现代视觉-语言-动作(VLA)策略预测整个动作块:在单次前向传播中发出一到两秒的协调运动。然而,动作块本质上是一个短多变量轨迹,但在这些模型中,它是由线性头解码的通用逐时间步隐藏标记序列。这未能充分利用两种运动结构。首先,频率:一个动作块叠加了平滑的全局趋势和跨时间尺度的精细修正运动,而单个标记将它们纠缠在一起。其次,跨相位几何:不同相位(到达、接触、抓取调整、稳定)的运动在表示空间中沿着非常不同、近似正交的方向展开,但对于任务而言却紧密相关,并沿时间轴出现。点积注意力通过内积对齐,因此它偏向于对齐的标记,在接近正交时最不敏感,使得网络需要通过迂回路径恢复这些关系。我们引入了时间-频率几何交叉注意力(TFGCA),一个即插即用模块,修复这两个盲点。TFGCA使用逐维可学习的平稳小波变换将动作块分解为时间-频率标记,每个时间标记通过交叉注意力从它们中检索信息,该交叉注意力通过可学习权重融合了点积(相似性)和楔积幅度(对近正交敏感)。零初始化残差在初始化时重现基础行为,因此可以将其添加到预训练的VLA上并联合微调。相对于同源基础模型,TFGCA在分布内LIBERO上平均提高+1.5,在分布外LIBERO-Plus上提高+6.3,在RoboTwin域随机化下的随机平均提高+28.5,在三个真实机器人AgiBot A2任务上的总体成功率提高+11.67个百分点,且在分布外增益更大。

英文摘要

Modern vision-language-action (VLA) policies predict a whole chunk of actions: one to two seconds of coordinated motion emitted in a single forward pass. Yet an action chunk is essentially a short multivariate trajectory, but inside these models it is a sequence of generic per-timestep hidden tokens decoded by a linear head. This under-serves two motion structures. First, frequency: a chunk superimposes a smooth global trend and fine corrective motion across time scales, and a single token entangles them. Second, cross-phase geometry: motions of different phases (reach, contact, grasp adjustment, settling) unfold along very different, near-orthogonal directions in representation space, yet are tightly related for the task and arise across the time axis. Dot-product attention scores alignment by an inner product, so it favors aligned tokens and is least sensitive near orthogonality, leaving such relationships for the network to recover through a detour. We introduce Time-Frequency Geometric Cross-Attention (TFGCA), a drop-in module repairing both blind spots. TFGCA uses a per-dimension learnable stationary wavelet transform to decompose the action chunk into time-frequency tokens, and each time token retrieves information from them via a cross-attention that fuses the dot product (similarity) with the wedge-product magnitude (sensitive to near-orthogonality) through a learnable weight. A zero-initialized residual reproduces the base behavior at initialization, so it can be dropped onto a pretrained VLA and fine-tuned jointly. Relative to the same-source base, TFGCA improves in-distribution LIBERO by +1.5 on average, the OOD LIBERO-Plus by +6.3, the randomized average under RoboTwin domain randomization by +28.5, and the overall success rate on three real-robot AgiBot A2 tasks by +11.67 points, with larger gains out of distribution.

↑