arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型中的可言语化表征形成全局工作空间

Verbalizable Representations Form a Global Workspace in Language Models

Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, Jack Lindsey

arXiv 2607.15495首次发表:更新:

发表机构

Anthropic(A÷)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨语言模型中可言语化表征,用雅可比透镜识别出J空间,其具有全局工作空间功能特性与结构特征,能揭示模型未显思考,发现训练后助手观点在工作空间,引入反事实反思训练,解码表征助于了解认知过程。

AI 中文摘要

人类大脑处理的所有信息中,只有一小部分能被有意识地获取,可用于言语报告、刻意控制和灵活推理。本文提出证据表明,大型语言模型中也出现了类似的功能区分。使用新的可解释性技术雅可比透镜,识别出模型在处理过程中随时准备言语化的表征,即J空间。J空间具有全局工作空间的功能特性,其内容可报告、可刻意调用和保持,用于无声推理的中间步骤等。J空间还具有与有意识获取相关的结构特征。在对齐审计中,它揭示了模型输出中未出现的战略思考等。研究发现训练后助手的观点会安装在工作空间中,还引入了反事实反思训练。这些结果表明语言模型维持着一组具有意识获取功能特征的特权表征,解码这些表征有助于了解正在进行的认知过程。

英文摘要

Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model's weights more widely than other representations. These properties make it a practical window into a model's unspoken thinking. In alignment audits, it reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model's outputs. We find that post-training installs the Assistant's point of view in the workspace, and we introduce counterfactual reflection training, which improves behavior by training only what a model would say if interrupted and asked to reflect. These results indicate that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑