arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaVSkip:跨层自适应视觉标记跳过以实现高效多模态大语言模型推理

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang

arXiv 2609.15131首次发表:更新:

发表机构

Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AdaVSkip提出跨层自适应视觉标记跳过方法,通过轻量级路由器动态决定跳过自注意力与MLP模块,结合两阶段训练,在LLaVA-NeXT-7B上减少53.2%计算量且保持性能。

AI 中文摘要

多模态大语言模型(MLLMs)需要大量计算来处理所有Transformer层中的众多视觉标记。大多数高效MLLM推理方法通过压缩视觉标记来利用水平冗余。除了标记减少之外,近期研究通过早期退出或固定层跳过利用垂直冗余。然而,我们发现这种冗余的程度和分布因输入而异,并且在自注意力模块和MLP模块之间有所不同。受这些观察的启发,我们提出了AdaVSkip,它为每一层配备两个轻量级路由器,独立决定视觉标记是通过还是跳过自注意力和MLP模块。这些决策共同定义了一条特定于输入的视觉计算路径,但其离散且不可微的性质使得学习有效路径具有挑战性。为解决这一挑战,我们开发了一个渐进式两阶段训练框架,仅更新路由器而保持骨干网络冻结。第一阶段通过监督训练建立初始路由策略,使用由模块级必要性得分得出的输入特定目标。为了进一步使路由策略与任务性能对齐,第二阶段使用强化学习,通过生成答案的直接反馈来优化路由决策。它将答案正确性奖励与跳过一致性奖励相结合,后者阻止过度保留视觉标记计算。在三个MLLM骨干网络上,AdaVSkip以显著更少的计算量保持了强大的任务性能。在LLaVA-NeXT-7B上,AdaVSkip将FLOPs减少了53.2%,同时保持了原始模型的平均性能。将其与视觉标记压缩相结合,这一减少幅度提高到91.2%,同时平均保留了原始性能的97.2%。

英文摘要

Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most methods for efficient MLLM inference exploit horizontal redundancy by compressing visual tokens. Beyond token reduction, recent studies exploit vertical redundancy through early exit or fixed-layer skipping. However, we find that the extent and distribution of this redundancy vary across inputs and differ between self-attention and MLP modules. Motivated by these observations, we propose AdaVSkip, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules. These decisions collectively define an input-specific visual-computation path, but their discrete and non-differentiable nature makes learning effective paths challenging. To address this challenge, we develop a progressive two-stage training framework that updates only the routers while keeping the backbone frozen. Stage I establishes an initial routing policy through supervised training with input-specific targets derived from module-wise necessity scores. To further align the routing policy with task performance, Stage II uses reinforcement learning to optimize routing decisions with direct feedback from generated answers. It combines an answer correctness reward with a skip-consistency reward that discourages excessive retention of visual-token computation. Across three MLLM backbones, AdaVSkip maintains strong task performance with substantially less computation. On LLaVA-NeXT-7B, AdaVSkip reduces FLOPs by 53.2\% while preserving the original model's average performance. Combining it with visual token compression increases this reduction to 91.2\%, while retaining 97.2\% of the original performance on average.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑