arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VLA-ACL:面向高效视觉-语言-动作模型的动作一致性视觉令牌剪枝

VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models

Owen Du, Yang Yue, Jie Zhang, Jiaqi Pi, Chi Bene Chen, Gao Huang

arXiv 2610.08133首次发表:更新:

发表机构

ETH Zurich; Tsinghua University(苏黎世联邦理工学院; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VLA-ACL通过动作一致性学习,在冻结基础模型的情况下剪枝视觉令牌,最高剪枝87.5%,计算减少75%,推理加速1.5倍,实现高效VLA模型。

AI 中文摘要

视觉-语言-动作(VLA)模型在机器人操作任务中表现出色,但在每个控制步骤中处理长令牌序列会产生高昂的计算成本,限制了实时部署。视觉令牌剪枝提供了一种直接解决方案,因为视觉补丁在输入序列中占主导地位且包含大量冗余。然而,现有方法要么依赖间接的免训练启发式方法(如注意力分数和运动阈值),要么需要对基础VLA模型进行昂贵的微调。我们提出VLA-ACL(动作一致性学习),该方法通过动作级监督学习一个轻量级视觉令牌剪枝策略,同时保持基础VLA模型完全冻结。训练目标鼓励从剪枝后的视觉上下文生成的动作与全上下文教师模型保持一致,并以真实动作作为辅助监督。这直接将令牌选择与其对下游控制输出的影响联系起来。在LIBERO和真实世界操作任务上的实验表明,VLA-ACL可剪枝高达87.5%的视觉令牌,同时保持竞争性性能,计算量减少高达75%,并实现1.5倍推理加速。这些结果比现有冻结VLA剪枝方法建立了更强的性能-效率权衡,并展示了动作级监督在视觉令牌选择中的价值。代码可在该https URL获取。

英文摘要

Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy. Existing approaches, however, either rely on indirect training-free heuristics, such as attention scores and motion thresholds, or require costly fine-tuning of the base VLA model. We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen. The training objective encourages actions produced from pruned visual contexts to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at https://github.com/du-owen/VLA-ACL.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑