发表机构
KAIST; NAVER AI Lab; NAVER(韩国科学技术院; NAVER人工智能实验室; NAVER公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究大型语言模型推理操作的内部几何结构,发现推理操作在中间层的表示中可分离,且操作对齐表示依赖上下文,证实语言模型的语言推理表达式与内部几何结构存在对应关系。
AI 中文摘要
大型语言模型的推理通过多种功能操作展开,例如问题构建、目标分解和演绎。尽管这些操作在文本中被明确区分,但人们对它们在表示空间中的几何组织方式知之甚少。为此,我们研究不同的推理操作是否在隐藏表示中表现出相应的几何结构。我们发现,这些操作在保留的表示中是可分离的,且可分离性在中间层达到峰值,并验证该结构无法用语义或位置混淆来解释。在各层中,逐标记的操作对齐在跨度上变得更分散,而相同的表面标记会根据其周围块的操作而被不同地表示。注意力掩码干预进一步表明,块起始处的操作对齐表示依赖于先前的推理上下文。因此,我们的工作表明,语言模型在语言推理表达式与其内部几何结构之间保持表示对应关系。代码和项目材料可在https URL获取。
英文摘要
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is known about how they are geometrically organized in representation spaces. To this end, we investigate whether distinct reasoning operations exhibit corresponding geometric structure in hidden representations. We find that operations are separable in held-out representations, with separability peaking in middle layers, and verify that this structure is not explained by lexical or positional confounds. Across layers, token-wise operation-alignment becomes more distributed over spans, while identical surface tokens are represented differently depending on the operation of its surrounding chunk. Attention-masking interventions further show that operation-aligned representations at chunk onset depend on preceding reasoning context. Consequently, our work demonstrates that language models maintain representational correspondence between linguistic reasoning expressions and their internal geometric structures. Code and project materials are available at https://github.com/naver-ai/beneath-cot.
CommentsTo appear in EMNLP 2026 Main Conference. 43 pages, 14 figures, 19 tables