ClosureBench:面向组合图推理的构造式基准
ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning
浏览论文内容
中文总结 AI 辅助
本研究推出ClosureBench基准,评估不同规模模型在组合图推理任务的表现,发现模型存在记忆效应、组合推理瓶颈,微调输出程序的模型可提升性能。
中文摘要 AI 辅助
我们推出ClosureBench,这是一个具有经编程验证的真实值的、面向组合图关系推理的构造式基准。与易受数据污染影响的固定测试集基准不同,ClosureBench可按需生成实例:每个任务的参考答案均通过在Ein张量逻辑语言中执行程序计算得出,确保机器可验证的正确性。该基准涵盖26个任务类别,分为三个组合级别(L1-L3),难度由三个独立维度控制:图大小、边密度和查询深度。我们评估了从15亿参数开源模型到前沿系统(o3、GPT-4.1、Gemini 2.5、Claude Sonnet 4)的模型,报告了三项发现。其一,由于该基准总能提供新实例,它可直接测量模型的记忆效应:在固定测试集上微调的模型,其在已见过实例与新实例上的准确率存在19.3个百分点的差距,这是静态测试集无法揭示的;我们将此范围限定为针对答案对的监督微调,而非预训练污染。其二,准确率随图大小和查询深度的增加而下降,且二者存在交互作用:模型会从图的自然语言描述中误读图,随后对错误的图进行正确推理,因此即使是最强的前沿模型,其在原子查询到组合查询上的表现也会下降。这一瓶颈是推理本身的属性,而非输入格式的属性:当图以JSON边列表或邻接矩阵而非 prose 形式给出时,该瓶颈依然存在。其三,微调后输出可执行程序而非答案的40亿参数模型,其在组合级别上的准确率几乎保持平稳,并以极低的令牌成本接近前沿模型的准确率(在保留实例上达94.3%);这一结果适用于Ein和Python+NetworkX两种程序目标,因此是经验证的程序合成的属性,而非某一种语言的属性。
英文摘要
Large language models fail on multi-step compositional reasoning, but measuring that failure is hard, because new models are trained on the benchmarks used to evaluate them. A fixed test set becomes a memorisation check soon after release. Constructive benchmarks avoid this by generating instances on demand. We introduce ClosureBench, a constructive benchmark for graph-relational logical reasoning. Each task is built from explicit primitives (reachability, degree, set operations, connectivity, aggregation), and its reference answer is computed by executing code that implements that logic exactly. Ground truth is therefore verified, and the supply of fresh instances is unlimited. The benchmark spans 26 task categories at three compositional levels, with three independent difficulty axes: graph size, edge density, and query depth. We evaluate models from 1.5B open weights to frontier systems (o3, GPT-4.1, Gemini 2.5, Claude Sonnet 4). Accuracy falls as graph size and query depth increase, and the two axes interact. The difficulty does not lie in the surface form, since it persists when the graph is given as a JSON edge list or an adjacency matrix rather than prose, nor in the reasoning rule, which models state correctly. It lies in carrying that rule out over the graph across many steps. A 4B model fine-tuned to emit verified programs instead of answers stays nearly flat across compositional levels, while every frontier model degrades. o3 falls from 96% on atomic queries to 82% on the most compositional; the 4B model holds at 93% at a fraction of the token cost. The program offloads multi-step execution to a runtime, and the model's remaining errors are almost entirely misread edges. Constructive generation also supports a direct memorisation check, comparing accuracy on seen and fresh instances.
发表机构
- AIM Research Lab(AIM研究实验室)
机构由 AI 辅助整理,请以论文原文为准。