arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

位置即你所需:一种用于基于多模态大语言模型的指代表达分割的免费午餐式令牌压缩策略

Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation

Yuhan Liu, Yixiong Zou, Yuhua Li, Ruixuan Li

arXiv 2608.26142首次发表:更新:

AI 中文总结

针对基于多模态大语言模型的指代表达分割任务的计算开销瓶颈,提出仅依赖位置信息的即插即用无训练令牌压缩方法PAYN,性能优于现有方法。

AI 中文摘要

指代表达分割(RES)旨在从复杂且隐含的文本查询中生成像素级分割掩码。尽管近期多模态大语言模型(MLLMs)的进展大幅提升了RES性能,但其过高的计算开销仍是关键瓶颈,而这一问题却鲜有研究。为填补该空白,我们首先在该任务上评估典型的令牌压缩方法,观察到令人惊讶的性能下降。本文旨在理解这一现象以寻求解决方案。通过大量实验,我们发现RES的令牌压缩需保留原始位置嵌入与局部邻域空间结构,表明视觉令牌位置信息比其他任务中更为关键。基于此见解,我们提出疑问:能否仅基于位置信息设计令牌压缩方法?为此,我们提出PAYN,一种即插即用、无需训练的令牌压缩方法,仅依赖位置信息。PAYN保留在每个局部邻域区域充分分布的令牌,同时严格保留原始位置索引,从而维持空间关系一致性。在多个RES基准上的实验表明,我们的方法优于现有令牌压缩方法,验证了在基于MLLM的RES任务中,位置确实是令牌压缩所需的全部。代码可在该https URL获取。

英文摘要

Referring Expression Segmentation (RES) aims to generate pixel-wise segmentation masks from complex and implicit textual queries. While recent advances in Multimodal Large Language Models (MLLMs) have substantially boosted RES performance, their prohibitive computational overhead remains a critical bottleneck, which, however, is rarely explored. To fill this gap, we first evaluate typical token compression methods on this task and observe a surprising performance degradation. In this paper, we aim to understand this phenomenon for a solution. By extensive experiments, we find that token compression for RES requires preserving the original position embeddings and local neighboring spatial structures, indicating that visual token position information is far more critical than in other tasks. Building on this insight, we ask: Can we design the token compression method purely based on the position information? Therefore, we propose PAYN, a plug-and-play, training-free token compression method that relies solely on position information. PAYN retains tokens that are adequately distributed in every local neighboring region while strictly preserving original positional indices, thereby maintaining spatial relational consistency. Experiments on multiple RES benchmarks demonstrate that our method outperforms existing token compression methods, verifying that position is indeed all you need for token compression in the MLLM-based RES task. Codes are avaliable at https://github.com/YuhanLiu231/PAYN.

CommentsAccepted by ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑