arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

线性表示是如何学习的?抽象动力学的精确解

How are linear representations learned? Exact solutions to the dynamics of abstraction

William W. Yang, Andrew M. Saxe, Peter E. Latham

arXiv 2607.08843首次发表:更新:

AI 中文总结

研究线性表示在训练中如何学习,开发框架研究训练时概念方向对齐即“抽象”过程,在最小线性网络获精确解,扩展到非线性网络分析其影响,证明衰减定律,为抽象提供动力学理论,用于改进大语言模型线性探针泛化。

AI 中文摘要

在人工和生物神经网络中,概念通常在表示空间中编码为一致的线性方向。在深度学习中,这一思想被称为线性表示假设,是许多基于线性探针的可解释性和控制方法的基础。然而,虽然之前的工作研究了训练后这些方向是否应该存在,但它们在训练期间出现的动力学仍然知之甚少。在此,我们开发了一个框架来研究训练期间概念方向的对齐——我们称之为“抽象”的过程。在最小线性网络设置中,我们获得了抽象完整轨迹的精确解。这些解揭示了控制抽象的关键分析原则。将理论扩展到非线性网络,我们分析了非线性选择如何影响抽象动力学。我们证明了一个显著的衰减定律。我们在开放模型中找到了该定律的证据,并将理论应用于改进大语言模型中的线性探针泛化。我们的结果为抽象提供了一个动力学理论,对可解释性和控制具有重要意义。

英文摘要

In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear probes, from concept detection to activation steering. Yet while prior work has studied whether such directions should exist $\textit{after}$ training, the dynamics of how they emerge $\textit{during}$ training remain poorly understood. Here, we develop a framework to study the alignment of concept directions during training - a process we call "abstraction". In a minimal linear network setting, we obtain exact solutions for the full trajectory of abstraction. These solutions reveal key analytic principles governing abstraction: (i) data and target geometry jointly determine abstraction at the end-of-learning, (ii) abstraction improves with network depth, and (iii) initialization scale controls the maximum abstraction reached during training. Extending our theory to nonlinear networks, we analyze how the choice of nonlinearity affects abstraction dynamics: erf networks approximate the linear theory, while abstraction in ReLU networks depends less on target geometry and more on input geometry. Across both, we prove a striking attenuation law: both nonlinearities weaken abstraction in activations relative to preactivations. We find evidence for this law in open models (DINOv3, Gemma 4) and apply our theory to improve linear probe generalization in LLMs. Together, our results provide a dynamical theory of abstraction with implications for interpretability and control.

Comments81 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑