跳过言语,重新聚焦视觉:多模态大语言模型中用于推理分割的潜在推理
Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models
浏览论文内容
中文总结 AI 辅助
提出LIRSeg,用潜在标记替代显式思维链进行推理分割,通过两阶段训练和三种信息机制,在多个基准上提升分割精度并大幅减少推理标记。
中文摘要 AI 辅助
推理分割旨在解释隐式文本查询并实现细粒度视觉感知,这对于人机交互和具身智能体等应用至关重要。现有方法通常在多模态大语言模型(MLLMs)定位目标之前生成显式的思维链(CoT)。尽管直观,这种显式言语推理会引入显著的注意力干扰:冗余的文本标记在感知标记生成期间干扰注意力,并增加视觉标记之间的有效距离。为解决此问题,我们提出LIRSeg,它用一组紧凑的可学习潜在标记完全替代显式CoT进行推理分割。LIRSeg分两个阶段训练:空间对齐将潜在标记锚定在对象相关的视觉证据上,GRPO进一步用分割奖励优化它们。为使这些紧凑的潜在标记更具信息量,我们从信息视角引入三种互补机制:用于选择信息丰富训练信号的极端优势采样、用于学习互补表征的解耦探索-稳定性更新,以及用于防止表征坍缩的潜在多样性增强。在基准上的大量实验表明,LIRSeg持续提高分割准确性和推理效率。与VisionReasoner基线相比,LIRSeg在ReasonSeg上实现4.9%的绝对gIoU提升,在MUSE上实现7.1%的提升,在MMR上实现4.7%的提升,同时实现约16倍的推理标记减少。代码可在补充材料中获得。
英文摘要
Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before localizing the target. Although intuitive, such explicit verbal reasoning introduces substantial attention interference: redundant textual tokens disrupt attention during perception-token generation and also increase the effective distance between visual tokens. To address this issue, we propose LIRSeg, which fully replaces explicit CoT with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages: spatial alignment grounds the latent tokens in object-relevant visual evidence, and GRPO further optimizes them with segmentation rewards. To make these compact latent tokens more informative, we introduce three complementary mechanisms from an information perspective: extreme-advantage sampling for selecting informative training signals, decoupled exploration-stability updates for learning complementary representations, and latent diversity amplification for preventing representational collapse. Extensive experiments on benchmarks demonstrate that LIRSeg consistently improves both segmentation accuracy and reasoning efficiency. Compared with the VisionReasoner baseline, LIRSeg achieves absolute gIoU improvements of 4.9% on ReasonSeg, 7.1% on MUSE, and 4.7% on MMR, while achieving a approximately 16x reduction in reasoning tokens. Code is available in supplementary materials.
发表机构
- National University of Defense Technology(国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。