arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

容量、响应性与对齐:什么使潜在结构具有可操作性

Capacity, Responsiveness and Alignment: What Makes a Latent Structure Actionable

Or Shafran, Mor Geva

arXiv 2610.06897首次发表:更新:

发表机构

Blavatnik School of Computer Science and AI, Tel Aviv University(特拉维夫大学布拉瓦特尼克计算机科学与人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将语言模型潜在结构的因果影响分解为容量、响应性和对齐三因素,发现三者均需高水平,并据此提出因果探针,在引导性能上提升17%-118%,概念检测仅降3%。

AI 中文摘要

在语言模型(LM)的激活空间中定位潜在结构,对于理解和控制其行为至关重要。然而,定位出的结构在因果影响上可能存在显著差异,这引出了一个问题:什么使一个结构具有可操作性。我们通过将因果影响分解为三个因素的乘积来应对这一问题,并实证表明它们作为可解释的、不同的约束发挥作用:容量,衡量模型输出对沿结构移动的敏感性;响应性,捕捉概念在当前上下文中的可提升程度;以及对齐,反映结构与概念在特定上下文中的表征的匹配程度。在4个LM家族和50个概念上,我们观察到因果有效性要求所有因素都处于高水平;低容量和低响应性分别将其降低84%和95%,而低对齐则可能逆转它,抑制概念表达。此外,我们发现因果性依赖于上下文,而非结构的内在属性,因果有效的方向形成一个随上下文变化的低维子空间。通过将线性探针的训练限制在该子空间上,我们引入了因果探针,其在跨模型的引导中实现了17%-118%的提升,而概念检测仅下降3%。

英文摘要

Localizing latent structures in the activation space of language models (LMs) is central to understanding and controlling their behavior. Yet, localized structures can differ substantially in their causal influence, raising the question of what makes a structure actionable. We tackle this question by casting causal influence as a product of three factors and showing empirically that they act as interpretable, distinct constraints: capacity, measuring the sensitivity of the model's output to movement along the structure, responsiveness, capturing how promotable the concept is given the current context, and alignment, reflecting how well the structure aligns with the context-specific representation of the concept. Across 4 LM families and 50 concepts, we observe that causal effectiveness requires all factors to be high; low capacity and responsiveness reduce it by 84% and 95%, respectively, while low alignment can reverse it, suppressing concept expression. Moreover, we find that causality is context-dependent rather than an intrinsic property of the structure, with causally effective directions forming a low-dimensional subspace that varies across contexts. By restricting the training of linear probes to this subspace, we introduce causal probes that achieve 17%-118% improvement in steering across models, with only 3% reduction in concept detection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑