发表机构
University of Southern California; National University of Singapore; Adobe Research; Texas A&M University; Rice University; The University of North Carolina at Chapel Hill(南加州大学; 新加坡国立大学; 奥多比研究院; 德克萨斯农工大学; 莱斯大学; 北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出JigShape拼图基准,发现零样本VLM大多缺乏几何推理能力,所有模型在大尺寸拼图上均出现性能崩塌,将可扩展几何推理确立为VLM的开放性挑战。
AI 中文摘要
拼图任务需要同时对视觉内容和几何约束进行推理,但现有基准使用矩形切割,在纹理重复区域会产生模糊的真实值。我们提出JigShape,这是一个带有榫卯互锁块的基准,几何约束提供了严格的局部兼容性要求,结合视觉内容可产生无歧义的真实值。该基准包含9.5万个实例,覆盖4×4到16×16四种网格密度。研究发现,零样本视觉语言模型(VLM)大多缺乏几何推理能力:5个前沿模型中仅1个(GPT-5.5)在4×4拼图上表现超过随机基线,其余均处于随机水平;监督微调在4×4拼图上准确率超过97%,但所有模型在更大网格上均出现性能崩塌:GPT-5.5在8×8拼图上准确率从70%降至接近随机,微调模型在12×12拼图上准确率甚至低于5%。这种“缩放悬崖”表明,现有架构无法在块数增加时保持一致的约束满足能力,JigShape将可扩展的几何推理确立为视觉语言模型的开放性挑战。
英文摘要
Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4$\times$4 to 16$\times$16), we find that \textbf{zero-shot VLMs largely lack geometric reasoning}: only one of five frontier models (GPT-5.5) exceeds random baseline on 4$\times$4 puzzles, while all others perform at chance level. While supervised fine-tuning achieves $>$97\% on 4$\times$4, \textbf{all models collapse on larger grids}: GPT-5.5 drops from 70\% to near-random on 8$\times$8, and even fine-tuned models fall below 5\% on 12$\times$12. This ``scaling cliff'' suggests current architectures cannot maintain consistent constraint satisfaction as the number of pieces increases. \ours{} establishes scalable geometric reasoning as an open challenge for vision-language models.