发表机构
Southern University of Science and Technology (SUSTech); National University of Singapore(南方科技大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TVCache,一种无训练框架,利用文本-视觉协同过滤注意力头并选择重用层,在匹配令牌保留率下提升VLA模型任务成功率,最高提升14.5个百分点并减少2.45倍计算量。
AI 中文摘要
视觉-语言-动作(VLA)模型能够实现泛化的机器人控制,但计算成本仍然很高。令牌缓存提供了一种无训练、即插即用的加速替代方案。然而,现有的VLA缓存并未充分利用VLA模型的一个关键归纳偏置:文本-视觉协同,即文本语义引导对任务相关区域的精确视觉定位。特别是,现有设计在注意力聚合中的头级可靠性和缓存重用中的层级稳定性方面考虑不足。为解决这一问题,我们提出了文本-视觉协同令牌缓存(TVCache),一种用于高效VLA推理的无训练框架。TVCache基于文本-视觉信息焦点过滤注意力头,以改善任务相关且物理一致的视觉定位。同时,我们引入了一种由文本-视觉熵差引导的重用层选择机制,以避免缓存不稳定的表示并改善缓存资源分配。在四个代表性VLA模型、两个仿真基准和真实世界机器人任务上的广泛实验证明了TVCache的有效性和泛化性。在匹配的令牌保留比率下,TVCache在可比计算成本下持续优于现有VLA缓存的任务成功率。在OpenVLA-OFT上,在12.5%保留率下,相较于VLA-Cache,平均成功率提高了最多14.5个百分点,同时相对于全令牌推理将FLOPs减少了2.45倍。
英文摘要
Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy, wherein textual semantics guide the precise visual grounding of task-relevant regions. In particular, existing designs insufficiently account for head-wise reliability in attention aggregation and layer-wise stability in cache reuse. To address this, we propose Text-Vision Synergistic Token Caching (TVCache), a training-free framework for efficient VLA inference. TVCache filters attention heads based on text-vision information focus to improve task-relevant and physically consistent visual grounding. Concurrently, we introduce a reuse-layer selection mechanism guided by text-vision entropy differences to avoid caching unstable representations and improve cache resource allocation. Extensive experiments across four representative VLA models, two simulation benchmarks, and real-world robotic tasks demonstrate the effectiveness and generality of TVCache. At matched token-retention ratios, TVCache consistently improves task success over existing VLA caching with comparable computational cost. On OpenVLA-OFT, it improves average success by up to 14.5 percentage points over VLA-Cache at 12.5% retention while reducing FLOPs by 2.45x relative to full-token inference.