发表机构
Hangzhou Institute for Advanced Study; University of Chinese Academy of Sciences; Computer Network Information Center; Chinese Academy of Sciences(杭州高等研究院; 中国科学院大学; 计算机网络信息中心; 中国科学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对统一多模态模型的系统级评估缺口,提出无标注的SGU语义闭环框架,通过感知-生成-推理流程评估模型综合能力,发现高性能模型存在自生成上下文推理局限,为下一代模型开发提供基准基础。
AI 中文摘要
随着大型视觉语言模型日益旨在将视觉生成与理解整合到单一参数空间中,以连贯方式评估这种结构的统一性仍是一项关键挑战。当前评估协议大多将生成能力与判别能力视为独立任务,导致统一多模态模型(UMMs)的系统级评估存在缺口。本研究提出了一种新颖的无标注评估框架——自生成理解(SGU),该框架通过语义闭环挑战探究统一模型的综合能力。无需新标注,SGU利用UMMs的理解与生成双重能力,要求其先感知图像并生成文本描述,随后基于该描述重构视觉上下文,最终对自生成输出进行推理。该流程提供了一种零成本测试平台,可生成专门用于将UMMs作为统一系统进行评估的综合性能分数。大量实验表明,即使是高性能的UMMs也常难以对自身生成的上下文进行推理,暴露出仅通过理解或生成的独立评估无法捕捉的局限性。本研究提供了一种互补的整体评估框架,并为下一代统一多模态模型开发的基准测试奠定了基础。
英文摘要
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs). In this work, we propose Self-Generative-Understanding (SGU), a novel, annotation-free evaluation framework that probes the integrated capabilities of unified models through a semantic closed-loop challenge. Without requiring new annotations, SGU leverages the dual understanding-and-generation abilities of UMMs by asking them to first perceive an image and produce a textual description, subsequently reconstruct a visual context based on that description, and finally perform reasoning over the self-generated output. This pipeline provides a zero-cost testbed that yields an integrated performance score specifically tailored for evaluating UMMs as unified systems. Extensive experiments show that even high-performing UMMs often struggle to reason over their own generated contexts, revealing limitations that are not captured by separate evaluations of understanding or generation alone. Our work provides a complementary holistic evaluation framework and offers a foundation for benchmarking the development of next-generation unified multimodal models.
Comments21 pages, 8 figures, 12 tables. Title updated; author contribution and corresponding-author notes added. Hao Zhang and Jiaxin Qi contributed equally; Jianqiang Huang is the corresponding author