面向视觉-语言-动作模型的自适应KV缓存复用的神经内部门控机制
Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models
- Graduate School of Information Science and Technology, The University of Tokyo(东京大学情报理工学系研究科)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对VLA模型视觉标记KV缓存复用未考虑模型不确定性的问题,提出轻量无训练的Gated VLA-Cache,通过监控动作标记对数几率间隔控制缓存,在LIBERO基准上恢复超100%精度并保留80%算力节省。
AI中文摘要:
视觉-语言-动作(VLA)模型通过单个自回归Transformer直接将相机图像和语言指令映射为运动指令。在实时控制中,它们仍会消耗大量算力,重新计算相邻帧间几乎无变化的视觉标记的键值(KV)表示。近期的VLA-Cache等工作通过复用视觉静态块的KV状态降低该成本,但其策略仅依赖观测空间启发式方法,未考虑模型自身的不确定性。我们提出Gated VLA-Cache,这是一种轻量、无需训练的扩展方法,将视觉相似性缓存与神经内省增强。该方法监控解码过程中可用的零成本置信信号——前两个预测动作标记的对数几率间隔,当间隔低于阈值时,使缓存失效并触发完整重新计算。在四个LIBERO基准套件上,使用OpenVLA和OpenVLA-OFT进行评估,Gated VLA-Cache在盲目缓存损害性能时提升了可靠性;在LIBERO-Goal和LIBERO-Long上,它恢复了超过100%的损失精度,同时保留了80%的算力节省。
英文摘要:
Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer. In real-time control, they still spend substantial compute recomputing key-value(KV) representations for visual tokens that barely change across neighboring frames. Recent work such as VLA-Cache reduces that cost by reusing KV states for visually static patches, but its policy relies only on observation-space heuristics and does not account for the model's own uncertainty. We propose Gated VLA-Cache, a lightweight, training-free extension that augments visual-similarity caching with neural introspection. The method monitors the logit margin between the top two predicted action tokens, a zero-cost confidence signal available during decoding. When the margin drops below a threshold, the cache is invalidated and a full recompute is triggered. Evaluated on four LIBERO benchmark suites with both OpenVLA and OpenVLA-OFT, Gated VLA-Cache improves reliability when blind caching hurts. On LIBERO-Goal and LIBERO-Long, it recovers over 100% of the lost accuracy while retaining 80% of the compute savings.