发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在自然文本中植入可控潜在变量,训练小型Transformer追踪其贝叶斯后验信念,发现其状态按马尔可夫链排列成环,为LLM信念状态与概念几何的关联提供了实证支持。
AI 中文摘要
大语言模型(LLM)被认为会追踪“信念状态”,即对支配语言的潜在变量的运行概率分布(Shai等人,2024;Sarfati等人,2026),但迄今为止,这仅在合成玩具数据和少量孤立案例研究中得到全面证明,且从未在经验层面与LLM特征的几何结构(可解释性在模型激活中发现的概念)建立关联。本研究在自然文本中植入可控潜在变量:LLM教师生成普通文本,同时我们在每个 token 处沿K=8个不相关的稀疏自编码器方向之一“潜意识”地引导它,活跃方向遵循环形马尔可夫链。在该语料库上训练的小型Transformer模型确实追踪了关于我们植入的潜在变量的贝叶斯后验信念,此外,它还将8种状态本身按马尔可夫链的精确顺序排列成环形,这为概念的几何结构可由其背后潜在变量的统计动力学形成提供了支持证据。
英文摘要
LLMs are thought to track "belief states," i.e., running probability distributions over the latent variables that govern language (Shai et al., 2024; Sarfati et al., 2026), but so far this has only been comprehensively demonstrated on toy synthetic data and in a few isolated case studies. It has also never been empirically connected to the geometry of LLM features (the concepts interpretability finds in model activations). In this work, we plant a controllable latent variable inside natural-looking text. An LLM teacher writes ordinary text while we "subliminally" steer it along one of K = 8 unrelated sparse autoencoder directions at each token, with the active directions following a ring-shaped Markov chain. A small transformer model trained on this corpus does indeed track the Bayesian posterior belief about our planted latent variable. Moreover, it also arranges the 8 states themselves on a ring, in the exact order of the Markov chain, which is supporting evidence that a concept's geometry can be formed by the statistical dynamics of the latent variable behind it.
Comments9 pages, 13 figures