arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HAP:基于跨模态对齐的头部自适应视觉Token剪枝

HAP: Head-Adaptive Visual Token Pruning via Cross-Modal Alignment

Yuanhao Sun, Huawei Ji, Yuan Jin, Cheng Deng, Luoyi Fu, Xinbing Wang

arXiv 2608.23921首次发表:更新:

发表机构

Shanghai Jiao Tong University; University of Edinburgh(上海交通大学; 爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉-语言模型预填充成本过高的问题,提出基于PAQ指标的头部自适应视觉Token剪枝方法,在18个基准上实现最优权衡,在LLaVA-1.5-7B上仅保留5.6% Token即可维持99.1%的原始性能,优于基线AutoPrune。

AI 中文摘要

近期的视觉-语言模型将高分辨率图像编码为长视觉Token序列,产生了高昂的预填充成本。为压缩这些序列,现有方法通过在所有头部中均匀平均文本到视觉的注意力来为每个视觉Token打分,该方法假设每个头部都与查询匹配。然而,我们的实证分析表明,未对齐的头部主导了平均值,放大了背景Token并淹没了细粒度线索。为解决这一问题,我们提出了PAQ(提示基础的注意力质量),这是一种量化每个头部将提示与图像区域对齐程度的指标。基于PAQ,我们的剪枝过程分为三个阶段:给定目标FLOPs预算,我们首先将Transformer层划分为组并为每组分配视觉Token预算;在每个组内,我们通过PAQ加权的Softmax将每个头部的注意力图聚合为组级矩阵;最后,我们通过该矩阵的幅值对视觉Token打分,并保留每组分配的预算。通过使用PAQ为头部加权,我们的方法通过更忠实地反映提示相关性的注意力信号对Token打分,而非通过均匀平均来稀释信号。在18个基准测试中,我们的方法实现了最优的权衡。具体而言,在LLaVA-1.5-7B(9个任务)上,仅保留5.6%的Token即可保留原始性能的99.1%,超过了最强的基线AutoPrune 4.2个百分点。代码可在该https URL获取。

英文摘要

Recent Vision-Language Models encode high-resolution images into long visual token sequences, incurring prohibitive prefill costs. To compress them, existing methods score each visual token by averaging text-to-visual attention uniformly across all heads, which assumes every head matches the query. However, our empirical analysis shows that misaligned heads dominate the average, amplifying background tokens and drowning out fine-grained cues. To address this, we propose PAQ (Prompt-Grounded Attention Quality), a metric quantifying how well each head aligns the prompt with image regions. Built on PAQ, our pruning proceeds in three stages. Given a target FLOPs budget, we first partition the transformer layers into groups and allocate a visual token budget to each. Within each group, we then aggregate per-head attention maps via PAQ-weighted softmax into a group-level matrix. Finally, we score visual tokens by this matrix's magnitude and retain the allocated budget per group. By weighting heads with PAQ, our method scores tokens by attention signals that more faithfully reflect prompt relevance, rather than diluting them through uniform averaging. Across 18 benchmarks, our method delivers state-of-the-art trade-offs. Specifically, on LLaVA-1.5-7B (9 tasks), retaining only \textbf{5.6\%} tokens preserves \textbf{99.1\%} of the original performance, surpassing the strongest baseline AutoPrune by 4.2 points. Code is available in https://github.com/baokou-fw2/HAP.

Journal refEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑