RoRA:面向角色的区域分配,用于多模态大语言模型中的视觉令牌剪枝
RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs
浏览论文内容
中文总结 AI 辅助
本研究针对多模态大语言模型视觉令牌剪枝的高开销问题,提出无训练框架RoRA,通过角色导向的区域分配实现高效剪枝,在多个模型系列上优于基线,大幅降低推理时间且保留高准确率。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)将图像编码为长视觉令牌序列,这使得预填充和键值缓存(KV-cache)存储的开销很大。现有的无训练剪枝方法通过重要性、多样性或空间覆盖度选择令牌,但将保留的令牌视为可互换的,且未明确跟踪哪些与对象相关的区域已被覆盖。我们提出RoRA,这是一个无训练框架,它将视觉令牌剪枝转化为面向角色的区域证据分配。在给定固定预算的情况下,RoRA将令牌划分为受保护的语义核心、补充上下文和细粒度细节。它首先通过位置先验和提示校准的对象先验来校准文本条件注意力,然后从高置信度锚点构建注意力锚定区域(AARs),作为已覆盖对象支持的轻量级代理。上下文主要在AARs之外探索,而少量由AAR引导的预算用于恢复局部细节;成对相似度仅用于上下文阶段的冗余过滤。在匹配的预算下,RoRA在LLaVA和Qwen-VL系列上始终优于强大的无训练基线,即使在激进的剪枝比例下也能保留大部分未剪枝的准确率,例如在LLaVA-1.5上进行88.9%的剪枝时,保留了96.5%的完整性能,并且在75%-90%的剪枝比例下,在Qwen3-VL上比D2Pruner提高了约5%。在66.7%的剪枝比例下,RoRA的令牌选择仅需0.7毫秒,端到端推理时间减少了24.6%,相当于在NVIDIA H800上比未剪枝的推理实现了1.33倍的加速。
英文摘要
Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing training-free pruning methods select tokens by importance, diversity, or spatial coverage, but treat retained tokens as interchangeable and do not explicitly track which object-related regions are already covered. We present RoRA, a training-free framework that casts visual token pruning as role-oriented regional evidence allocation. Given a fixed budget, RoRA partitions tokens into a protected semantic core, complementary context, and fine-grained detail. It first calibrates text-conditioned attention with a positional prior and a prompt-calibrated object prior, then builds Attention-Anchored Regions (AARs) from high-confidence anchors as lightweight proxies for covered object support. Context is explored mainly outside AARs, while a small AAR-guided budget restores local detail; pairwise similarity is used only for context-stage redundancy filtering. Under matched budgets, RoRA consistently outperforms strong training-free baselines across LLaVA and Qwen-VL families, retaining most of the unpruned accuracy even at aggressive pruning ratios, e.g., 96.5% of full performance at 88.9% pruning on LLaVA-1.5, and improving over D2Pruner by about 5% on Qwen3-VL at 75-90% pruning. At a 66.7% pruning ratio, RoRA requires only 0.7 ms for token selection and reduces end-to-end inference time by 24.6%, corresponding to a 1.33x speedup over unpruned inference on an NVIDIA H800.