arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38081cs.LGq-bio.NCstat.ML

穿越神经网络解空间:Hessian零空间延拓

Traversing the solution space of neural networks with Hessian Null Space Continuation

Ann Huang, Mitchell Ostrow, Zhouyang Lu, William T. Redman, Leo Kozachkov, Kanaka Rajan

首次发表
浏览论文内容

中文总结 AI 辅助

提出Hessian零空间延拓(HNC)方法,利用局部曲率遍历神经网络解空间,首次证明局部模式连通区域内存在多种内部机制,并揭示表示多样性,为模型合并与编辑提供新途径。

中文摘要 AI 辅助

在单一任务上,深度网络可以学习到许多解,具体取决于其优化器、训练数据、架构和超参数。其中许多解是模式连通的:它们并非权重空间中的孤立点,而是由低损失区域连接在一起。然而,这些区域内网络的内部计算如何变化尚不清楚。另一条平行研究方向揭示了神经表示的简并性:许多网络以不同的内部结构达到相似的训练损失。但是,这些解在权重空间中如何关联仍不明确。我们统一了这两个子领域,并首次证明在权重空间的局部模式连通区域内存在多种不同的内部机制。为此,我们提出了Hessian零空间延拓(HNC),这是一种可扩展的方法,利用局部曲率穿越保持网络功能的权重空间区域,并可引导至具有指定性质的解。在记忆任务上训练的RNN中,HNC在保持行为的同时达到了截然不同的表示和动力学。在ImageNet训练的Vision Transformer中,HNC找到的表示与原始网络的差异大于任何独立训练的具有不同架构或目标的模型。在强化学习智能体中,HNC在相当回报下揭示了一种不同的导航策略,并在AI安全Gridworld中暴露了奖励黑客行为。最后,HNC测量了解集的局部几何结构,展示了模型大小和任务复杂度如何塑造其维度和功能敏感性。我们的结果表明,在单个训练解附近存在令人惊讶的大量表示多样性,而标准基于梯度的优化无法看到这些多样性。HNC识别并量化了这种多样性,为解空间的机制理解以及模型合并、编辑和微调开辟了新的可能性。

英文摘要

On a single task, deep networks can learn many solutions, depending on their optimizer, training data, architecture, and hyperparameters. Many of these solutions are mode-connected: rather than isolated points in weight space, they are connected by low-loss regions. Yet how their internal computation varies within these regions is unknown. A parallel line of work has identified the degeneracy of neural representations: many networks reach similar training loss with distinct internal structures. However, it is unclear how these solutions are related in weight space. We unify these subfields and show for the first time that many different internal mechanisms exist within a local mode-connected region in weight space. To do so, we introduce Hessian Null Space Continuation (HNC), a scalable method that uses local curvature to traverse regions of weight space that preserve network function, and can be steered toward solutions with specified properties. In RNNs trained on a memory task, HNC reaches drastically different representations and dynamics with maintained behavior. In ImageNet-trained Vision Transformers, HNC finds representations that differ more from the original network than any independently trained model with a different architecture or objective. In reinforcement-learning agents, HNC uncovers a distinct navigation strategy at comparable return and exposes reward hacking in an AI Safety Gridworld. Finally, HNC measures the local geometry of the solution set, showing how model size and task complexity shape its dimension and functional sensitivity. Our results show that a surprisingly large amount of representational diversity exists near a single trained solution, unseen by standard gradient-based optimization. HNC identifies and quantifies this diversity, opening new possibilities for mechanistic understanding of solution spaces and for model merging, editing, and fine-tuning.

发表机构

  • Harvard University(哈佛大学)
  • Harvard Medical School(哈佛医学院)
  • Kempner Institute(肯普纳研究所)
  • Massachusetts Institute of Technology(麻省理工学院)
  • Brown University(布朗大学)
  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑