arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AnchorPrune:用于视觉令牌剪枝的相关性锚定上下文扩展

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning

Kyuan Oh, Bumsoo Kim

arXiv 2607.07033首次发表:更新:

发表机构

Chung-Ang University(崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大型视觉语言模型推理成本高、视觉令牌冗余问题,提出无需训练的AnchorPrune框架,通过构建相关性锚点并扩展,自适应确定锚点大小,分配预算恢复上下文,提升了模型的准确率与效率,尤其在严重压缩下效果显著。

AI 中文摘要

大型视觉语言模型推理成本高昂,因为高分辨率输入会引入数千个视觉令牌,其中许多对给定查询是冗余的。现有剪枝方法常结合查询相关性和令牌多样性,但在激进压缩下目标会冲突。我们引入AnchorPrune,一个无需训练的框架,先构建受保护的相关性锚点,再用互补视觉上下文扩展它。它从相关性排序的令牌新颖性概况自适应确定锚点大小,保留紧凑的查询关键证据集,并通过重要性加权新颖性分配剩余预算以恢复相对于锚点的信息丰富、非冗余上下文。这种有序设计防止上下文扩展取代不可或缺的查询线索,同时提高整体视觉覆盖率。AnchorPrune轻量级、架构感知,无需重新训练或修改模型。在图像和视频视觉语言模型及基准测试中,它始终优于无需训练的基线方法,尤其在严重压缩下。在LLaVA-NeXT-7B上,AnchorPrune仅使用2880个视觉令牌中的160个就保留了97.6%的全令牌性能。这些结果确立了相关性锚定上下文扩展作为高效多模态推理的有效原则。

英文摘要

Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are redundant for a given query. Existing pruning methods often combine query relevance and token diversity, yet these objectives can conflict under aggressive compression: relevance-driven selection may overconcentrate the budget on correlated local evidence, while diversity-driven selection may suppress indispensable tokens or retain distinct but uninformative regions. We introduce AnchorPrune, a training-free framework that first constructs a protected relevance anchor and then expands it with complementary visual context. AnchorPrune adaptively determines the anchor size from the novelty profile of relevance-ranked tokens, preserving a compact set of query-critical evidence, and allocates the remaining budget through importance-weighted novelty to recover informative, non-redundant context relative to the anchor. This ordered design prevents contextual expansion from displacing indispensable query cues while improving overall visual coverage. AnchorPrune is lightweight, architecture-aware, and requires neither retraining nor model modification. Across image and video vision-language models and benchmarks, it consistently improves the accuracy-efficiency trade-off over training-free baselines, particularly under severe compression. On LLaVA-NeXT-7B, AnchorPrune preserves 97.6% of full-token performance using only 160 of 2,880 visual tokens. These results establish relevance-anchored contextual expansion as an effective principle for efficient multimodal inference. Code is available at https://github.com/MULTI-cau/AnchorPrune.

CommentsECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑