AI 中文总结
该研究推出PolyComp基准测试集,考察多模态模型的组合式3D空间推理能力,测试了GPT-5.6 Sol等模型的准确率与成本,并发布了120个问题。
AI 中文摘要
我们推出PolyComp,这是一个经程序生成和验证的基准测试集,重点考察视觉识别与组合式空间推理能力。在每个问题中,模型必须从四个选项中识别出哪一对多立方体(polycube)组件能够组合形成目标实体。该基准测试集包含四个几何类别共120个问题,每个问题有三种不同的呈现格式,分别为单张图像或多张图像。随机猜测的基线准确率为25%。在三种呈现格式下(每个模型共呈现360个问题),全力运行的GPT-5.6 Sol达到50.0%的准确率(95%问题簇置信区间为43.3-56.7%),每个呈现问题的平均成本为0.951美元;全力运行的Claude Fable 5达到39.4%的准确率(33.1-46.1%),每个呈现问题的成本为0.701美元;高思考等级的Gemini 3.1 Pro Preview达到27.5%的准确率(22.8-32.5%),接近25%的随机猜测基线,每个呈现问题的成本为0.350美元。观察到的跨几何类别的准确率差异大于跨呈现格式的差异。我们提出了问题开发与评估协议、成本与token核算方法,并发布了这120个问题。
英文摘要
We introduce PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning. In each problem, a model must identify which of four options shows a pair of polycube components that can be combined to form a target solid. The benchmark contains 120 problems across four geometry families, and each problem has three different presentation formats using either a single image or multiple images. The random guessing baseline is 25%. Across the three presentations (360 presented problems per model), GPT-5.6 Sol with max effort attains 50.0% accuracy (95% problem-cluster CI 43.3-56.7%) at a mean cost of \$0.951 per presented problem, Claude Fable 5 with max effort attains 39.4% (33.1-46.1%) at \$0.701, and Gemini 3.1 Pro Preview with thinking level high attains 27.5% (22.8-32.5%), near the 25% random guessing baseline, at \$0.350. The observed accuracy spread across geometry families is larger than across presentation formats. We present a problem development and evaluation protocol, cost and token accounting, and release the 120 problems.