arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经动力学作为量化单元的组合

Neural Dynamics as the Composition of Quantized Units

Jacopo Minniti, Aravinth Kulanthaivelu, Richard Sproat

arXiv 2609.32487首次发表:更新:

AI 中文总结

本文提出量子抽象,将训练视为量化单元的有序获取,推导其动力学,并在组合任务和Transformer实验中验证,揭示宏观行为与微观计算的连接。

AI 中文摘要

深度学习通常在两个层面上被解释:宏观层面,通过缩放定律所总结的损失中的总体趋势;微观层面,通过神经元、特征和电路。一个核心挑战是理解这些层面如何连接,以便我们能够解释基本计算如何组合并共同塑造宏观行为。为此,我们研究了一种中间抽象,其中训练被描述为量子的有序获取:可重复使用的计算突然获得,并在示例间以二进制方式激活以减少损失。通过近似群体梯度更新,我们推导出量子的获取动力学。这产生了一个由需求(计算在示例间被需要的频率)和条件复杂度(在已有计算可用的情况下获取该计算的难度)所决定的获取优先级。在一个布尔组合任务中,我们推导出获取顺序的预测,并展示了错开的离散获取如何产生平滑的总体损失,并且在量子组合的某些几何结构下,产生缩放定律。然后,我们训练一个Transformer将数字映射为英文数字名称,并从其检查点轨迹中恢复候选量子。从这些单元中,我们构建了一个模型,该模型保留了Transformer的大部分行为,同时暴露了与理论一致的可解释的潜在计算和获取动力学。此外,量子结构可以作为训练目标以提高Transformer的泛化能力。这些结果共同表明,量子抽象可以为研究各种宏观现象提供有用的计算原子。

英文摘要

Deep learning is commonly interpreted at two levels: the macroscopic, through aggregate trends in loss summarized by scaling laws, and the microscopic, through neurons, features, and circuits. A central challenge is understanding how these levels connect, so that we can explain how elementary computations compose and collectively shape macroscopic behavior. To this end, we study an intermediate abstraction in which training is described as the ordered acquisition of quanta: reusable computations acquired suddenly and binary-activated across examples to reduce loss. By approximating population-gradient updates, we derive quanta's acquisition dynamics. This yields an acquisition priority governed by demand, how frequently a computation is required across examples, and conditional complexity, how difficult that computation is to acquire given those already available. In a Boolean compositional task, we derive predictions for acquisition order and show how staggered discrete acquisitions can produce smooth aggregate loss and, under certain geometries of quanta composition, give rise to scaling laws. We then train a Transformer to map numerals to English number names and recover candidate quanta from its checkpoint trajectory. From these units, we construct a model that preserves much of the Transformer's behavior while exposing interpretable latent computations and acquisition dynamics consistent with the theory. Separately, the quanta structure can serve as training targets to improve transformer generalization. Together, these results suggest the quanta abstraction can provide useful computational atoms for studying a variety of macroscopic phenomena.

Comments22 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑