发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究预训练语言模型从隐马尔可夫模型预测下一个观测值的能力背后算法,通过三阶段流程,先实证比较缩小候选算法范围,再推导理论联系并验证,最后引入主激活探针揭示低维线性表示及变化,将其上下文行为与内部机制联系起来。
AI 中文摘要
大语言模型(LLMs)通过上下文学习(ICL)从隐马尔可夫模型(HMMs)预测下一个观测值的能力显著,但该能力背后的算法仍未确定。先前工作提出了几个候选算法但未达成共识,且都未基于模型内部激活。本文通过三阶段流程弥补差距。首先,将LLM行为与一系列候选算法进行实证比较,缩小到三类。其次,推导这三类之间的理论联系并展示如何在上下文中由Transformer实现,在小型训练Transformer中验证。最后,引入主激活探针(PAP),揭示低维线性表示驱动模型预测并跟踪ICL性能,还展示了这些表示如何随HMM机制属性变化。结果将预训练LLMs的上下文行为与潜在内部机制联系起来,推进了对LLMs在HMMs上执行ICL的理解。
英文摘要
Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has proposed several candidates without consensus, and none has been grounded in the model's internal activations. We close this gap with a three-stage pipeline. First, we empirically compare LLM behavior against a suite of candidate algorithms and narrow the space to three classes -- though no single class explains LLM behavior across all HMM settings and sequence lengths. Second, we derive theoretical connections between the three classes and show how each can be implemented in-context by a Transformer, validating the construction in a small trained Transformer. Third, returning to pre-trained LLMs, we introduce the Principal Activations Probe (PAP), a layer-wise probing and intervention method that isolates algorithmic signals in model activations. PAP reveals low-dimensional linear representations that causally drive model predictions and track empirical ICL performance. PAP further reveals how these representations shift with properties of the underlying HMM regime; distinct computational stages are localized to different layers. Together, our results connect the in-context behavior of pre-trained LLMs to the underlying internal mechanisms and advance our understanding of how LLMs perform ICL on HMMs.