arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

统一多模态模型中视觉生成何时能助力视觉理解?

When Does Visual Generation Help Visual Understanding in Unified Multimodal Models?

Yubo Zhu, Zhehan Kan, Jingyi Yang, Miaolin Chen, Jinbo Xing, Kai Zhu, Zijian Wang, Sheng Zhong, Wei Tong

arXiv 2608.22174首次发表:更新:

发表机构

Nanjing University; Tsinghua University; Fudan University(南京大学; 清华大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出VGAU-Diag评估框架,发现统一多模态模型的视觉生成仅在简单任务上助益理解,瓶颈多在理解侧,有效生成需瞄准理解瓶颈。

AI 中文摘要

统一多模态模型(UMMs)可同时执行理解与生成任务,由此引出核心问题:视觉生成能否提升理解能力?现有评估提供的证据相互矛盾,但混淆了任务难度、推理范式以及生成与理解间的闭环交互。我们提出VGAU-Diag,一种用于视觉生成辅助理解的细粒度评估框架。该框架按难度分层样本,支持对多种推理范式的统一评估,并采用“先知辅助参考协议”。我们的分析显示,生成的视觉辅助在较简单实例上有帮助,但随着推理复杂度提升会变得不可靠。先知辅助诊断进一步揭示,主要瓶颈常在于视觉理解侧而非视觉生成侧,因为当前UMMs难以利用即使是忠实的视觉辅助。我们还表明,有效视觉生成应瞄准视觉理解瓶颈,而非增加更多推理步骤,并识别出从任务无关噪声、误导性似是而非的指导到最终有用辅助的三阶段转变。这些发现将有助于指导更好的模型开发,代码可在指定链接获取。

英文摘要

Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understanding? Existing evaluations provide mixed evidence, but confound task difficulty, reasoning paradigms, and the closed-loop interaction between generation and understanding. We introduce VGAU-Diag, a fine-grained evaluation framework for vision generation-assisted understanding. It stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Assisted Reference Protocols. Our analysis shows that generated visual aids help on easier instances but become unreliable as reasoning complexity increases. Oracle-assisted diagnosis further reveals that the main bottleneck often lies on the visual-understanding side rather than the visual-generation side, as current UMMs struggle to leverage even faithful visual aids. We also show that effective visual generation should target visual-understanding bottlenecks rather than add more reasoning steps, and identify a three-stage transition from task-irrelevant noise, to misleading plausible guidance, and finally to useful assistance. These findings would be useful to guide the development of better UMMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑