arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17376cs.LGcs.CL

大型语言模型在上下文中发展出信念状态几何结构

Large Language Models Develop Belief State Geometry In-Context

Daniel Balcells, Andrew Jun Lee, Chirag Rastogi, Paul M. Riechers, Adam Shai, Xavier Poncini

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过探测隐马尔可夫模型数据提示下的开源LLMs,发现其上下文学习中的信念状态可线性解码,干预该子空间可保持预测质量,表明ICL近似最优贝叶斯预测。

中文摘要 AI 辅助

基于下一词元预测训练的大型语言模型(LLMs)展现出显著的上下文学习(ICL)能力,然而支撑ICL的表示仍未被充分理解。我们在一个受控环境中研究此类表示:用从隐马尔可夫模型(HMMs)生成的数据提示LLMs,并探测相应的信念状态——即给定观测到的词元历史后,对HMM隐藏状态的后验分布。在六个开源LLMs中,使用从40个具有非平凡信念结构的HMMs中选取的数据进行提示,我们发现信念状态可以从残差流激活中线性解码,在HMM和LLM组合中,峰值探针$R^2$值在0.83至0.99之间,分布范围从早期层到后期层。为确立功能相关性,我们通过补丁和引导直接干预探针识别的子空间,导致下游预测质量与未篡改模型相当,而对照组则显著降低性能。综合来看,这些结果提供了表示层面的证据,表明开源LLMs中的ICL近似于对上下文推断的生成模型的最优贝叶斯预测。更广泛地说,我们的发现扩展了先前将输入分布结构与激活几何联系起来的结果:从明确在HMM数据上训练的玩具网络到生产规模的LLMs。

英文摘要

Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from hidden Markov models (HMMs) and probing for the corresponding belief state -- the posterior distribution over the HMM's hidden states given the observed token history. Across six open-source LLMs prompted with data from 40 HMMs selected for non-trivial belief structure, we find that belief states are linearly decodable from residual stream activations, with peak probe $R^2$-values from 0.83-0.99 across HMM and LLM combinations, ranging from early to late layers. To establish functional relevance, we intervene directly on the probe-identified subspace via patching and steering, resulting in downstream prediction quality on the order of the untampered model, while controls degrade performance substantially. Together, these results provide representation-level evidence that ICL in open-source LLMs approximates optimal Bayesian prediction over a context-inferred generative model. More broadly, our findings extend prior results linking input-distribution structure to activation geometry: from toy networks trained explicitly on HMM data to production-scale LLMs.

发表机构

  • UCLA(加州大学洛杉矶分校)
  • UIUC(伊利诺伊大学厄巴纳-香槟分校)
  • Astera Institute(阿斯特拉研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑