arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JigShape:通过拼图任务评估视觉语言模型的视觉几何推理能力

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Shawn Li, Wei Yang, Jike Zhong, Jiate Li, Jiawei Yang, You Qin, Ryan Rossi, Franck Dernoncourt, Roger Zimmermann, Yue Wang, Zhengzhong Tu, Vicente Ordonez, Mohit Bansal, Yue Zhao

arXiv 2607.27670首次发表:更新:

发表机构

University of Southern California; National University of Singapore; Adobe Research; Texas A&M University; Rice University; The University of North Carolina at Chapel Hill(南加州大学; 新加坡国立大学; 奥多比研究院; 德克萨斯农工大学; 莱斯大学; 北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出JigShape拼图基准,发现零样本VLM大多缺乏几何推理能力,所有模型在大尺寸拼图上均出现性能崩塌,将可扩展几何推理确立为VLM的开放性挑战。

AI 中文摘要

拼图任务需要同时对视觉内容和几何约束进行推理,但现有基准使用矩形切割,在纹理重复区域会产生模糊的真实值。我们提出JigShape,这是一个带有榫卯互锁块的基准,几何约束提供了严格的局部兼容性要求,结合视觉内容可产生无歧义的真实值。该基准包含9.5万个实例,覆盖4×4到16×16四种网格密度。研究发现,零样本视觉语言模型(VLM)大多缺乏几何推理能力:5个前沿模型中仅1个(GPT-5.5)在4×4拼图上表现超过随机基线,其余均处于随机水平;监督微调在4×4拼图上准确率超过97%,但所有模型在更大网格上均出现性能崩塌:GPT-5.5在8×8拼图上准确率从70%降至接近随机,微调模型在12×12拼图上准确率甚至低于5%。这种“缩放悬崖”表明,现有架构无法在块数增加时保持一致的约束满足能力,JigShape将可扩展的几何推理确立为视觉语言模型的开放性挑战。

英文摘要

Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4$\times$4 to 16$\times$16), we find that \textbf{zero-shot VLMs largely lack geometric reasoning}: only one of five frontier models (GPT-5.5) exceeds random baseline on 4$\times$4 puzzles, while all others perform at chance level. While supervised fine-tuning achieves $>$97\% on 4$\times$4, \textbf{all models collapse on larger grids}: GPT-5.5 drops from 70\% to near-random on 8$\times$8, and even fine-tuned models fall below 5\% on 12$\times$12. This ``scaling cliff'' suggests current architectures cannot maintain consistent constraint satisfaction as the number of pieces increases. \ours{} establishes scalable geometric reasoning as an open challenge for vision-language models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑