发表机构
Central South University; National University of Singapore(中南大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GenPuzzle基准测试通过12个轨道2005个视觉谜题评估图像生成模型的推理能力,最强模型仅得40.57分,揭示其在逻辑和几何方面的不足。
AI 中文摘要
最近的图像生成系统日益结合多模态理解、推理和合成,这表明它们可能不仅仅是渲染合理的场景。然而,现有的评估强调美学、提示对齐、组合性或基于文本的答案,使得这些系统能否解决视觉问题并忠实地在像素中表达解决方案尚不明确。我们引入了GenPuzzle,一个以推理为中心的图像生成基准。GenPuzzle包含12个轨道中的2,005个问题,涵盖图案补全、空间构造、迷宫、数独、填字游戏、七巧板、棋盘游戏、火柴棍谜题、正投影和数学视觉证明。每个任务提供一个视觉谜题,并要求输出一个图像,该图像在保持输入状态的同时执行逻辑上有效的解决方案。GenPuzzle使用特定于任务的评估协议:离散网格输出被转录并以编程方式验证,而视觉复杂的输出则使用分层、多维或二进制的多模态大语言模型(MLLM)评分标准进行评估。我们进一步通过测量与人类参考分数的一致性来选择自动评判者。在三个前沿生成器中,最强的模型仅达到40.57的宏观总体分数,揭示了在逻辑、几何、状态保持和指令执行方面的频繁失败。GenPuzzle提供了一个从图像渲染向视觉问题解决进展的测试平台。
英文摘要
Recent image generation systems increasingly combine multimodal understanding, reasoning, and synthesis, suggesting that they may do more than render plausible scenes. Yet existing evaluations emphasize aesthetics, prompt alignment, compositionality, or text-based answers, leaving unclear whether these systems can solve visual problems and faithfully express solutions in pixels. We introduce GenPuzzle, a benchmark for reasoning-centric image generation. GenPuzzle contains 2,005 problems across 12 tracks, spanning pattern completion, spatial construction, mazes, Sudoku, nonograms, tangrams, board games, matchstick puzzles, orthographic projection, and mathematical visual proof. Each task provides a visual puzzle and requires an image output that preserves the input state while executing a logically valid solution. GenPuzzle uses task-specific evaluation protocols: discrete grid outputs are transcribed and verified programmatically, while visually complex outputs are assessed with tiered, multidimensional, or binary multimodal large language model (MLLM) rubrics. We further select the automatic judge by measuring agreement with human reference scores. Across three frontier generators, the strongest model reaches only 40.57 Macro Overall, revealing frequent failures in logic, geometry, state preservation, and instruction execution. GenPuzzle provides a testbed for measuring progress from image rendering toward visual problem solving.