arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07088cs.CVcs.AI

RoRA:面向角色的区域分配,用于多模态大语言模型中的视觉令牌剪枝

RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs

Qiyanhui Lu, Han Wu, Rongjian Xu, Tingzhang Luo, Cheng Fan, Xinghao Chen, Minjing Dong, Jufeng Yang, Jianyuan Guo

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对多模态大语言模型视觉令牌剪枝的高开销问题,提出无训练框架RoRA,通过角色导向的区域分配实现高效剪枝,在多个模型系列上优于基线,大幅降低推理时间且保留高准确率。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)将图像编码为长视觉令牌序列,这使得预填充和键值缓存(KV-cache)存储的开销很大。现有的无训练剪枝方法通过重要性、多样性或空间覆盖度选择令牌,但将保留的令牌视为可互换的,且未明确跟踪哪些与对象相关的区域已被覆盖。我们提出RoRA,这是一个无训练框架,它将视觉令牌剪枝转化为面向角色的区域证据分配。在给定固定预算的情况下,RoRA将令牌划分为受保护的语义核心、补充上下文和细粒度细节。它首先通过位置先验和提示校准的对象先验来校准文本条件注意力,然后从高置信度锚点构建注意力锚定区域(AARs),作为已覆盖对象支持的轻量级代理。上下文主要在AARs之外探索,而少量由AAR引导的预算用于恢复局部细节;成对相似度仅用于上下文阶段的冗余过滤。在匹配的预算下,RoRA在LLaVA和Qwen-VL系列上始终优于强大的无训练基线,即使在激进的剪枝比例下也能保留大部分未剪枝的准确率,例如在LLaVA-1.5上进行88.9%的剪枝时,保留了96.5%的完整性能,并且在75%-90%的剪枝比例下,在Qwen3-VL上比D2Pruner提高了约5%。在66.7%的剪枝比例下,RoRA的令牌选择仅需0.7毫秒,端到端推理时间减少了24.6%,相当于在NVIDIA H800上比未剪枝的推理实现了1.33倍的加速。

英文摘要

Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing training-free pruning methods select tokens by importance, diversity, or spatial coverage, but treat retained tokens as interchangeable and do not explicitly track which object-related regions are already covered. We present RoRA, a training-free framework that casts visual token pruning as role-oriented regional evidence allocation. Given a fixed budget, RoRA partitions tokens into a protected semantic core, complementary context, and fine-grained detail. It first calibrates text-conditioned attention with a positional prior and a prompt-calibrated object prior, then builds Attention-Anchored Regions (AARs) from high-confidence anchors as lightweight proxies for covered object support. Context is explored mainly outside AARs, while a small AAR-guided budget restores local detail; pairwise similarity is used only for context-stage redundancy filtering. Under matched budgets, RoRA consistently outperforms strong training-free baselines across LLaVA and Qwen-VL families, retaining most of the unpruned accuracy even at aggressive pruning ratios, e.g., 96.5% of full performance at 88.9% pruning on LLaVA-1.5, and improving over D2Pruner by about 5% on Qwen3-VL at 75-90% pruning. At a 66.7% pruning ratio, RoRA requires only 0.7 ms for token selection and reduces end-to-end inference time by 24.6%, corresponding to a 1.33x speedup over unpruned inference on an NVIDIA H800.

补充信息

↑