arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GEB-Bench:多视角呈现的抽象结构

GEB-Bench: Abstract Structures Told in Many Voices

Tong Zhang, Zhiyuan Shi, Yun Peng, Tao Xie

arXiv 2608.04111首次发表:更新:

发表机构

Fudan University; Peking University(复旦大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出GEB-Bench基准,评估12种模型发现其在结构识别与跨视角映射间存在差距,仅前沿模型能缩小该差距,且表面复杂性会增加模型负担。

AI 中文摘要

模型能否在看到河流三角洲和闪电时,识别出它们共享同一结构?我们推出GEB-Bench,这是一个以抽象结构 motif(主题)为单元的基准,其灵感源于《哥德尔、艾舍尔、巴赫:集异璧之大成》(Gödel, Escher, Bach),这些结构包括自指、怪圈、莫比乌斯扭转。每个 motif 以多种视角呈现:其结构构成的自然场景、通过可机械校验的形式装置来演绎该结构的民间故事、数学定理以及程序框架;表面参数被视为干扰变量,不纳入评分。motif、视角以及它们之间的结构变化构成了一个小型跨模态类别,GEB-Bench 的任务即围绕该类别展开。我们评估了12种开源及闭源模型,发现抽象失败具有规律性。核心发现是识别与跨视角映射之间存在差距:模型在单一视角内识别结构的表现远优于跨视角传递结构;所有模型均存在该差距,只有前沿级别的模型才具备足够的映射能力来缩小该差距。两种模式支持这一结论:错误与设计的形式几何的契合度,强于与实测感知几何的契合度,且不同厂商的前沿模型会得出相同的错误答案;表面复杂性会对所有读取结构的模型造成负担,模型容量仅能提供缓冲而非免疫。GEB-Bench 完全可生成,且附带其 pipeline(流程)发布。

英文摘要

Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstract structural motif--self-reference, a strange loop, a Mobius twist--in the spirit of Godel, Escher, Bach. Each motif is told in several voices: a natural scene whose composition is the structure, a folk story whose telling enacts it through a mechanically checkable form device, a mathematical theorem, and a programmatic skeleton; surface parameters are declared nuisance variables and never scored. Motifs, voices, and the structural changes between them form a small cross-modal category, and GEB-Bench's tasks are its questions. Evaluating twelve open and proprietary models, we find that abstraction failure is lawful. The central finding is a gap between recognition and cross-voice mapping: models identify a structure within one voice far better than they carry it across voices; every model pays this tax, and mapping strong enough to narrow it appears only at the frontier tier. Two patterns support it. Errors align more strongly with the designed formal geometry than with measured perceptual geometries, and frontier models from different vendors converge on the same wrong answers; and surface complexity taxes every model that reads structure, with capacity buying headroom rather than immunity. GEB-Bench is fully generative and released with its pipeline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑