发表机构
Nanjing University; Tsinghua University; Fudan University(南京大学; 清华大学; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出VGAU-Diag评估框架,发现统一多模态模型的视觉生成仅在简单任务上助益理解,瓶颈多在理解侧,有效生成需瞄准理解瓶颈。
AI 中文摘要
统一多模态模型(UMMs)可同时执行理解与生成任务,由此引出核心问题:视觉生成能否提升理解能力?现有评估提供的证据相互矛盾,但混淆了任务难度、推理范式以及生成与理解间的闭环交互。我们提出VGAU-Diag,一种用于视觉生成辅助理解的细粒度评估框架。该框架按难度分层样本,支持对多种推理范式的统一评估,并采用“先知辅助参考协议”。我们的分析显示,生成的视觉辅助在较简单实例上有帮助,但随着推理复杂度提升会变得不可靠。先知辅助诊断进一步揭示,主要瓶颈常在于视觉理解侧而非视觉生成侧,因为当前UMMs难以利用即使是忠实的视觉辅助。我们还表明,有效视觉生成应瞄准视觉理解瓶颈,而非增加更多推理步骤,并识别出从任务无关噪声、误导性似是而非的指导到最终有用辅助的三阶段转变。这些发现将有助于指导更好的模型开发,代码可在指定链接获取。
英文摘要
Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understanding? Existing evaluations provide mixed evidence, but confound task difficulty, reasoning paradigms, and the closed-loop interaction between generation and understanding. We introduce VGAU-Diag, a fine-grained evaluation framework for vision generation-assisted understanding. It stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Assisted Reference Protocols. Our analysis shows that generated visual aids help on easier instances but become unreliable as reasoning complexity increases. Oracle-assisted diagnosis further reveals that the main bottleneck often lies on the visual-understanding side rather than the visual-generation side, as current UMMs struggle to leverage even faithful visual aids. We also show that effective visual generation should target visual-understanding bottlenecks rather than add more reasoning steps, and identify a three-stage transition from task-irrelevant noise, to misleading plausible guidance, and finally to useful assistance. These findings would be useful to guide the development of better UMMs.