arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型激活中的语义搜索特征

Signatures of semantic search in the activations of large language models

Luke Leckie, Peter M. Todd, Jacob G. Foster

arXiv 2609.35599首次发表:更新:

发表机构

Indiana University; Santa Fe Institute(印第安纳大学; 圣塔菲研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过机制可解释性技术,证明大型语言模型在语义流畅性任务中,其内部激活状态存在与人类类似的探索与利用的语义搜索特征,并能通过转向向量调控切换行为。

AI 中文摘要

在语义流畅性任务(SFT)中回忆概念列表(例如动物)时,人类和大型语言模型(LLMs)都会将其输出组织为相关项目(例如海洋动物)的聚类,这些聚类之间穿插着策略性的切换。在人类中,这种模式可以通过语义觅食过程来解释,其中不同的神经和行为特征伴随着聚类内产生(“利用”)和聚类间切换(“探索”)。LLMs是否同样在其内部状态中表征这两种搜索模式尚不清楚。在此,我们应用一系列机制可解释性技术来为此提供证据。在研究1中,我们使用雅可比透镜(J-lens),该透镜将中间层残差流表示映射到词元级激活,以表明概念级激活预测切换。首先,我们发现切换与低下一词元激活同时发生。此外,随着最强J-lens激活集合(J空间)中当前正在产生的类别的项目逐渐耗尽,切换概率上升,类似于补丁觅食中的探索-利用决策。然后,我们表明,抽象类别相关标签(例如“水”)的中间层J-lens激活在预期切换到该类别时增加。我们通过推导针对类别切换的转向向量,确认这些表示对切换具有因果影响。在研究2中,我们识别出在切换期间及预期切换时被激活的通用残差流方向。通过沿这些方向转向激活,我们偏向于增加或减少切换率。我们的研究将语义觅食框架扩展到人工智能,并提供证据表明LLMs在口头表达概念信息时,维持着探索和利用的独特表征特征。

英文摘要

When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise their output into clusters of related items (e.g., sea animals) that are punctuated by strategic switches between clusters. In humans, this pattern can be explained by a semantic foraging process, whereby distinct neural and behavioural signatures accompany within-cluster production ("exploit") and between-cluster switching ("explore"). Whether LLMs likewise represent these two search regimes within their internal states is unknown. Here, we apply a range of mechanistic interpretability techniques to provide evidence for this. In Study 1, we use the Jacobian lens (J-lens), which maps intermediate-layer residual-stream representations to token-level activations, to show that concept-level activations predict switching. First, we find that switching coincides with low next-token activations. Moreover, the probability of switching rises as the set of strongest J-lens activations (the J-space) becomes depleted of items from the category currently being produced, analogous to explore-exploit decision-making during patch foraging. We then show that middle-layer J-lens activations of abstract category-related labels (e.g., "water") increase in anticipation of switching into that category. We confirm these representations to causally influence switching by deriving steering vectors that target category switching. In Study 2, we identify generic residual stream directions that are activated during and in anticipation of switching. By steering activations along these directions, we bias increased or decreased rates of switching. Our study extends the semantic foraging framework to artificial intelligences and provides evidence that LLMs maintain distinct representational signatures for exploration and exploitation as they verbalise conceptual information.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑