arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自回归马赛克:探测仅文本语言模型的二维空间推理能力

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

Ashwin Nedungadi, Stefan Oehmcke, Stefan Lüdtke

arXiv 2608.30751首次发表:更新:

发表机构

Institute for Visual & Analytic Computing (VAC), University of Rostock(罗斯托克大学视觉与分析计算研究所(VAC))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究引入AM-Bench基准,发现仅文本LLMs的二维空间表现取决于模型和输出介质,无法仅用代码生成能力解释,其开放式布局表现存在显著差异。

AI 中文摘要

仅在文本和代码上训练的大型语言模型(LLMs)有时能生成绘制出可识别图像的程序,但目前尚不清楚这是反映了模型对二维空间布局的内部表征,还是仅具备将空间描述转化为代码的能力。我们引入了自回归马赛克(AM-Bench)这一基准,用于分离上述影响因素:首先是翻译任务,为模型提供以文字形式完整指定的图片几何结构作为提示,要求模型生成能绘制该图片的代码;其次是布局任务,要求模型根据未完全指定的提示合成图像。针对8个仅文本和代码的开放权重模型开展的实验显示,所有模型都能可靠地将指定几何结构转化为代码,但它们的开放式布局表现存在显著差异,说明这些差异无法仅用代码生成能力来解释。输出介质消融实验进一步表明,模型使用的表达接口或介质很重要:将过程代码替换为原始SVG可提升所有模型的布局得分。最后,对模型激活的探测显示,生成前就存在粗略的布局计划,但该计划仅反映提示所隐含的布局;生成过程中,模型会跟踪不断演变的几何状态,而非执行最初固定的计划。总体而言,这些结果表明,仅文本LLMs的二维空间表现取决于模型本身和输出介质,且无法仅用代码生成能力来解释。

英文摘要

Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or simply the ability to translate spatial descriptions into code. We introduce Autoregressive Mosaics (AM-Bench), a benchmark that separates these factors: First, a translation task gives a model a fully specified geometry of a picture in words as a prompt and asks for the code that produces it. Second, a layout task requires the model to compose an image from an underspecified prompt. Across eight open-weight text-and-code-only models, all models reliably translate specified geometry into code, but their open-ended layout performance differs substantially, indicating that these differences are not explained by code-generation ability alone. An output-medium ablation further shows that the interface or medium of expression that the model uses matters: replacing procedural code with raw SVG improves layout scores across all models. Finally, probing model activations shows that a coarse layout plan is present before generation, but reflects only the layout implied by the prompt. During generation, models track the evolving geometric state instead of executing an initially fixed plan. Overall, these results show that 2D spatial performance in text-only LLMs depends on both the model and the output medium, and is not explained by code-generation ability alone.

CommentsPre-Print

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑