arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21883cs.CV

VIG:作为多模态思维链压缩奖励信号的视觉信息增益

VIG: Visual Information Gain as a Reward Signal for Multimodal Chain-of-Thought Compression

Wen Luo, Xiaohan Yi, Xiaotao Huang, Liqun Huang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出VIG奖励机制,通过提升视觉信息密度优化多模态CoT的准确率与效率权衡,无需额外资源,在多类基准及不同规模模型上均有效。

中文摘要 AI 辅助

多模态大推理模型通常依赖冗长的思维链(CoT)轨迹,其中大量标记(如重复的视觉描述、自我反思及其他与视觉无关的填充内容)会增加推理成本,却对答案生成无贡献。现有CoT压缩方法仅优化输出长度,从未衡量推理标记是否实际基于图像。我们提出VIG(Visual Information Gain,视觉信息增益),这是一种基于信息论的GRPO奖励,通过图像降低每个推理标记预测不确定性的程度来对其评分。VIG通过同一策略的两次前向传播在线计算,一次使用图像,一次不使用图像,因此无需参考链、外部注释或辅助奖励模型。在六个主要多模态推理基准、三种Qwen3-VL-Thinking模型规模(2B/4B/8B),以及额外的8B规模R1-Onevision-Bench评估中,VIG始终提升了准确率与效率的权衡,支持我们的核心主张:高效多模态推理源于提升视觉信息密度,即每个推理标记都通过锚定图像来确立自身价值,而非通过设定长度预算。我们的源代码可在该https URL获取。

英文摘要

Multimodal large reasoning models often rely on long Chain-of-Thought (CoT) traces in which a substantial fraction of tokens, such as repeated visual descriptions, self-reflection, and other visually-disengaged filler, inflate inference cost without contributing to the answer. Existing CoT compression methods optimize output length but never measure whether a reasoning token is actually grounded in the image. We propose \textbf{VIG} (Visual Information Gain), an information-theoretic GRPO reward that scores each reasoning token by how much the image reduces its predictive uncertainty. VIG is computed online from two forward passes of the same policy, one with and one without the image, so no reference chains, external annotations, or auxiliary reward models are needed. Across six main multimodal reasoning benchmarks and three Qwen3-VL-Thinking model sizes (2B/4B/8B), plus an additional R1-Onevision-Bench evaluation on 8B, VIG consistently improves the accuracy--efficiency trade-off, supporting our central claim: \emph{efficient multimodal reasoning emerges from raising visual information density, where every reasoning token earns its place by anchoring to the image, rather than from imposing a length budget.} Our source code is available at https://github.com/chaser682/vig.

发表机构

  • School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑