发表机构
University of Illinois Urbana-Champaign; New York University(伊利诺伊大学厄巴纳-香槟分校; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有生成模型存在的推理-生成差距问题,提出包含2000个样本的RIG-BENCH基准,评估四类认知任务下的推理驱动图像生成,为下一代逻辑型生成模型开发提供诊断框架。
AI 中文摘要
近期统一生成模型(UGMs)与世界模拟器的进展在视觉感知与合成领域取得了前所未有的成果,但这些模型主要依赖表层事件对齐,对高级视觉推理能力的探索不足。真正的视觉生成智能需要“推理到生成”的能力,即从视觉输入中推断潜在规则,并通过精确、逻辑约束的视觉结果呈现解决方案。本文提出RIG-BENCH,这是一个新颖的综合基准,在概念类、变换类、模式与结构类、场景类四个认知要求较高的领域系统评估推理驱动图像生成(RIG),包含2000个精心筛选的样本,是对RIG的严格压力测试。本文对当前最先进的UGMs及图像/视频生成模型的广泛评估显示,存在显著的推理-生成差距,模型常生成局部合理但全局不合逻辑的输出。RIG-BENCH为指导下一代基于逻辑的UGMs与世界模拟器的开发提供了重要的诊断框架。
英文摘要
Recent advancements in unified generative models (UGMs) and world simulators have achieved unprecedented results in visual perception and synthesis. However, these models primarily rely on surface-level event alignment, leaving the capacity for high-level visual reasoning underexplored. True visual generative intelligence demands "Reasoning-to-Generation", an ability to infer latent rules from visual inputs and manifest solutions through precise, logically constrained visual outcomes. We introduce RIG-BENCH, a novel comprehensive benchmark that systematically evaluates Reasoning-driven Image Generation (RIG) across four cognitively demanding domains: Concept-based, Transformation-based, Pattern & Structure, and Scenario-based. Featuring 2000 curated samples, RIG-BENCH serves as a rigorous stress test for RIG. Our extensive evaluations of state-of-the-art UGMs and image/video generation models reveal a significant reasoning-generation gap, wherein models frequently produce locally plausible but globally illogical outputs. RIG-BENCH provides a vital diagnostic framework to guide the development of next-generation, logically grounded UGMs and world simulators.