arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21247cs.CVcs.RO

视觉-语言-动作模型中用于令牌压缩的刚可察觉差异建模

Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models

Zhuoyuan Li, Rui Zhao, Jin Wang, Hanwei Zhu, Cong Zhang, Giuseppe Valenzise, Weisi Lin, Kin-Man Lam

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对视觉-语言-动作模型的令牌压缩问题,提出Action-JND方法,通过动作容忍度准则优化压缩,在LIBERO基准实验中提升了激进压缩下的压缩可靠性。

中文摘要 AI 辅助

令牌压缩已成为降低大型基础模型推理成本的关键技术,令牌剪枝、KV缓存复用等方法已在视觉-语言模型中广泛应用,近来也被应用于具身智能体。在具身智能体中,令牌不仅支持感知与语义理解,还直接影响对延迟敏感的闭环机器人动作预测。现有方案通常利用冗余度或重要性线索(如视觉相似度、注意力分数、显著性)指导压缩,但这些线索仅间接衡量安全压缩的关键因素:令牌发生多大改变才会导致下游动作出现不可接受的偏差。这种依赖接收方的容忍度与刚可察觉差异(JND)原理密切相关:经典JND描述人类视觉系统的信号容忍度,面向机器的JND则将该概念扩展至下游机器响应。基于此,本文提出Action-JND,将JND建模扩展至具身感知,通过视觉-语言-动作(VLA)策略在闭环控制中的语言条件动作响应定义可察觉性,仅当令牌变化引发的动作偏差在容忍范围内时,该变化才被视为可接受。为实现这一概念,我们在深度视觉特征空间开发轻量级逐令牌JND估计器,以预测在保留策略响应的同时最大可容忍的扰动,所得动作容忍度分数可作为即插即用准则,用于VLA压缩范式(包括过时KV复用和令牌剪枝),优先选择动作容忍度低的令牌进行压缩。在LIBERO基准上,使用OpenVLA和OpenVLA-OFT开展的实验表明,Action-JND在激进压缩率下可持续提升压缩可靠性。

英文摘要

Token compression has become a key technique for reducing the inference cost of large foundation models, with approaches such as token pruning and KV-cache reuse widely adopted in vision-language models and recently explored for embodied agents. In embodied agents, tokens not only support perception and semantic understanding but also directly affect latency-sensitive closed-loop robot action prediction. Existing schemes typically guide compression using redundancy or importance cues, such as visual similarity, attention scores, and saliency. However, these cues only indirectly measure the key factor for safe compression: how much a token can change before causing an unacceptable deviation in downstream actions. This receiver-dependent tolerance is closely related to the principle of just noticeable difference (JND). Classical JND characterizes signal tolerance in the human visual system, while machine-oriented JND extends this concept to downstream machine responses. Building on this progression, we introduce Action-JND, which extends JND modeling to embodied perception by defining noticeability through the language-conditioned action response of a vision-language-action (VLA) policy in closed-loop control. A token change is considered admissible only when the induced action deviation remains within a tolerated margin. To realize this concept, we develop a lightweight token-wise JND estimator in deep visual-feature space to predict the maximum tolerable perturbation while preserving policy responses. The resulting action-tolerance score serves as a plug-and-play criterion for VLA compression paradigms, including stale-KV reuse and token pruning, prioritizing action-tolerant tokens for compression. Experiments on the LIBERO benchmark with OpenVLA and OpenVLA-OFT demonstrate that Action-JND consistently improves compression reliability, especially under aggressive compression ratios.

发表机构

  • The Hong Kong Polytechnic University(香港理工大学)
  • Nanyang Technological University(南洋理工大学)
  • Bytedance Inc.(字节跳动公司)
  • Shandong University(山东大学)
  • CNRS(法国国家科学研究中心)
  • CentraleSupelec(中央理工学院)
  • Université Paris-Saclay(巴黎萨克雷大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑