arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经网络优化中的信息分配动态

The Anatomy of Implicit Bias: Information Allocation in Neural Network Training

Zhang Gongyue, Wang Zhiyong, Liu Donghan, Ren Weihong, Sheng Yixuan, Liu Honghai

arXiv 2607.07156首次发表:更新:

发表机构

State Key Laboratory of Robotics and Systems, Harbin Institute of Technology Shenzhen(机器人系统国家重点实验室(哈尔滨工业大学深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该论文从信息分配动态角度出发,将优化器隐含偏差解释为权重与偏差类参数路径间训练信号的相对分配,通过连续预处理指数\(p\)描述调整,分析其在训练中的形成机制,揭示此相对更新分配对参数轨迹和泛化行为的影响。

AI 中文摘要

不同的优化器有不同的更新偏差,但这些偏差通常是隐含的。现有研究主要从最终解的几何角度分析或控制此类偏差。然而,优化器偏差在训练过程中如何形成仍缺乏明确的内部机制。本文提出了信息分配动态的观点,将优化器隐含偏差解释为权重类和偏差类参数路径之间训练信号的相对分配,这种分配可由连续的预处理指数\(p\)描述和调整。为表征此机制,首先在最小线性模型中分析权重和偏差对同一残差信号的更新贡献。权重校正项保留输入相关的残差信号,偏差校正项保留残差均值方向,它们对应于残差信号的不同投影路径。将预处理更新代入残差更新方程后,优化器可通过不同的预处理因子改变权重校正项和偏差校正项的相对强度。因此,优化器隐含偏差不仅体现在最终解或全局训练轨迹中,还体现在不同参数路径上训练信号的相对写入率中。总体而言,本文将优化器隐含偏差的分析从解空间几何转移到训练期间的更新动态,揭示了权重类和偏差类参数之间的相对更新分配是影响参数轨迹和泛化行为的重要动态机制。

英文摘要

Implicit bias is usually explained as the preference of an optimization process for certain final solutions and their geometry. This view helps explain where a model finally stops. It gives less direct explanation of how this bias is formed during training. This paper proposes a training-time information allocation view. Under this view, optimization forms a writing pattern for error signals across parameter paths, coordinate channels, and sample regions. This paper builds a set of observable allocation diagnostics. These diagnostics include gradient demand, actual update injection, coordinate gain induced by exponential moving averages, channel-level update ratios, and sample-wise loss distributions. To separate training progress from internal allocation, this paper introduces a collapse--persistence analysis. Under matched training loss, if external loss statistics collapse but internal allocation ratios remain separated, then the factor changes the internal allocation of the training signal. Overall, this paper extends the analysis of implicit bias from final-solution geometry to training-time signal allocation. The main claim is that implicit bias is not only reflected by the final solution. It is also reflected by which parameter paths, coordinate channels, and sample regions receive the error signal first and more strongly during training. Based on this view, this paper places different training factors into a unified information-allocation diagnostic framework. The framework gives a mechanism-level explanation of training-time implicit bias. It also provides a basis for future optimization methods that control training progress and signal allocation separately.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑