arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36967cs.RO

超越令牌重要性:保留空间支架以实现高效的视觉-语言-动作推理

Beyond Token Importance: Preserving Spatial Scaffolds for Efficient Vision-Language-Action Inference

Jiayu Chen, Shuyong Gao, Jingkai Jia, Xiaosheng Bu, Jiyuan Fu, Lingyi Hong, Kaixun Jiang, Yipan Xu, Wenqiang Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA模型剪枝忽视空间信息的问题,提出GeoScaffold方法,通过空间分区与最远点采样保留空间支架,仅用20%视觉令牌保持93.2%成功率并实现1.78倍加速。

中文摘要 AI 辅助

现有的视觉-语言-动作(VLA)模型剪枝策略主要根据任务级语义相关性选择单个视觉令牌,而忽略了机器人操作所需的空间信息。为了检验这一局限性,我们构建了一个简单的Stride基线,该基线沿展平的一维视觉序列均匀采样令牌,代表一种纯粹的几何剪枝策略。令人惊讶的是,Stride在某些剪枝比率下优于语义剪枝和随机剪枝,但当令牌预算仅略微减少时,其性能会崩溃。我们通过空间覆盖半径来表征这一现象,该半径定义为剪枝后保留令牌集所引发的最大空间盲区。我们的分析揭示了保留令牌的空间结构与任务成功之间存在强相关性,表明可靠的VLA剪枝不仅需要保留与任务相关的令牌,还需要保留场景的空间支架。受此诊断启发,我们提出了GeoScaffold,一种无需训练的视觉令牌剪枝方法,该方法将每个图像划分为空间区域,使用任务相关性权重分配区域间令牌预算,并通过最远点采样选择区域内支架令牌以减小局部覆盖半径。在pi 0.5和LIBERO上,GeoScaffold仅保留20%的视觉令牌,同时保持了93.2%的平均成功率,并且相对于未剪枝基线实现了1.78倍的预填充加速。

英文摘要

Existing VLA pruning strategies primarily select individual visual tokens according to task-level semantic relevance, while overlooking the spatial information required for robotic manipulation. To examine this limitation, we construct a simple Stride baseline that uniformly samples tokens along the flattened one-dimensional visual sequence, representing a purely geometric pruning strategy. Surprisingly, Stride outperforms semantic pruning and random pruning at certain pruning ratios, but collapses when the token budget is only slightly reduced. We characterize this phenomenon through the spatial coverage radius, defined as the largest spatial blind spot induced by the retained token set after pruning. Our analysis reveals a strong correlation between the spatial structure of retained tokens and task success, suggesting that reliable VLA pruning requires preserving not only task-relevant tokens but also the spatial scaffold of the scene. Motivated by this diagnosis, we propose GeoScaffold, a training-free visual token pruning method that partitions each image into spatial regions, allocates inter-region token budgets using task-relevance weights, and selects intra-region scaffold tokens via farthest point sampling to reduce the local coverage radius. On pi 0.5 and LIBERO, GeoScaffold retains only 20% of visual tokens while preserving a 93.2% average success rate, and achieves a 1.78 times prefill speedup over the unpruned baseline.

发表机构

  • Fudan University(复旦大学)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑