arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文关系的几何结构:语言模型按提及顺序寻址事实

The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention

Yufa Zhou

arXiv 2610.00910首次发表:更新:

发表机构

Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出序数寻址假设,发现大语言模型通过事实在上下文中的提及顺序而非名称来寻址事实,并验证了事实地址的按序、可引导、低秩和涌现性等特性。

AI 中文摘要

人类推理依赖于命题中对象之间的相互关系。关系如何组织上下文内容的语言表示?我们给大语言模型(LLM)一个上下文中的事实列表(例如,爱丽丝吃苹果。鲍勃吃梨。),并测量当问题从爱丽丝吃什么切换到鲍勃吃什么时,其隐藏状态如何变化。对许多列表取平均后,这种变化是一个引导向量,我们称之为“序数向量”。它通过事实的“提及顺序”(即事实在上下文中陈述的顺序)来指向一个事实。我们发现,LLM 通过问题的提及顺序来表示问题所询问的事实,而不是通过问题中包含的名称。我们将此表述为“序数寻址假设”:每个提及顺序在模型状态中都有一个“事实地址”,该地址在所有上下文中共享,问题将状态移动到其所询问事实的事实地址,而上下文则提供该事实的内容。在 Qwen、Gemma 和 Llama 中,事实地址具有以下特点:(1) 按提及顺序排列:查询状态按事实的顺序组织,而非名称的顺序,即使一个事实有多个主语也是如此;(2) 可引导:将序数向量添加到关于新列表中第一个事实的问题上,会使模型用该列表的第二个事实来回答;(3) 低秩:它们张成一个低秩子空间,其中最先提及的事实最容易到达,这与人类记忆惊人地相似;(4) 涌现性:它们在后期中间层共享,适用于 1.5B 到 32B 参数,并在预训练早期形成。语言模型通过事实被提及的位置来访问已陈述的事实,这加深了我们对 LLM 推理的理解。

英文摘要

Human reasoning depends on how objects are related within propositions. \textit{How do relations organize the language representations of contextual contents?} We give an LLM a list of facts in its context (e.g., \emph{Alice eats an apple. Bob eats a pear.}) and measure how its hidden state changes when the question switches from what Alice eats to what Bob eats. Averaged over many lists, this change is a steering vector, which we call the \emph{ordinal vector}. It points to a fact by its \emph{order of mention}, the order in which the facts were stated in the context. We find that LLMs represent the fact a question asks about by its order of mention, not by the name the question contains. We state this as the \textit{ordinal addressing hypothesis}: each order of mention has a \emph{fact address} in the model's state, shared by all contexts, and a question moves the state to the fact address of the fact it asks about, while the context supplies what that fact says. Across Qwen, Gemma, and Llama, fact addresses are (1) \emph{ordered by mention}: query states are organized by the order of facts, not of names, even when one fact has multiple subjects; (2) \emph{steerable}: added to a question about the first fact of a new list, the ordinal vector makes the model answer with the second fact of that list; (3) \emph{low-rank}: they span a low-rank subspace in which the first-mentioned fact is the easiest to reach, surprisingly similar to human recall; and (4) \emph{emergent}: they are shared in late-middle layers, hold from 1.5B to 32B parameters, and form early in pretraining. Language models reach a stated fact by where it was mentioned, deepening our understanding of LLM reasoning.

CommentsCode: https://github.com/MasterZhou1/order-of-mention

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑