arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

数据可预测性塑造Transformer训练中的Weibull权重尺度增长规律

Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training

Tiexin Ding

arXiv 2608.23573首次发表:更新:

AI 中文总结

该研究发现训练前计算的二元条件熵可预测Transformer训练中Weibull权重尺度的增长规律,构建了学习率相关的量化关系,在两种架构上验证了规律的普适性。

AI 中文摘要

训练后的Transformer权重幅值可由两参数Weibull分布描述,其形状参数k≈1.2,在各层及不同模型间保持稳定,因此尺度参数λ承载了训练引发的大部分权重变化。何种语料属性决定了λ的增长幅度?利用训练前可计算的无训练统计量——二元条件熵D=H(下一个|前一个),在受控的语料损坏族范围内,我们发现了一个学习率η相关的规律:λ² - λ₀² = C₀(η) + C₁(η)(Hᵣ - D)^0.59,其中Hᵣ是匹配预算的打乱基线。该凸指数源自独立测量的数据侧饱和关系,而非直接拟合增长曲线。扣除两个随η变化的系数后,23次学习率跨一个数量级的运行结果均以单位斜率坍缩至(Hᵣ - D)^0.59(R²=0.941;直接按η拟合的效果更弱,R²≈0.82)。由于D在训练前即可计算,该规律是一种正向预测器:端到端自验证能以5.7%的相对误差复现保留的同族权重增长。该规律在模型及层级分辨率上均成立,在两种测试架构间函数形式保持不变,仅系数变化;它还标记了边界:跨语料预测会高估代码数据,表明冗余性是更广泛的Φ(D,R,A,H)数据-权重框架的第二个维度。

英文摘要

A trained transformer's weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape $k \approx 1.2$ is stable across layers and models, so the scale $λ$ carries most training-induced movement. What corpus property sets how much $λ$ grows? Using the bigram conditional entropy $D = H(\text{next} \mid \text{prev})$, a training-free statistic computed before training, we find across controlled corruption families a learning-rate-conditioned law, $λ^2 - λ_0^2 = C_0(η) + C_1(η)(H_r - D)^{0.59}$, where $H_r$ is a matched-budget shuffle baseline. The convex exponent is inherited from an independently measured data-side saturation relation rather than fitted directly to the growth curve. After removing the two per-$η$ coefficients, 23 runs spanning an order of magnitude in learning rate collapse onto $(H_r - D)^{0.59}$ with unit slope ($R^2 = 0.941$; direct per-$η$ fits are weaker, $R^2 \approx 0.82$). Because $D$ is computed before training, the law is a forward predictor: an end-to-end self-validation recovers held-out within-family weight growth with 5.7% relative error. The readout holds at model and per-layer resolutions and across two tested architectures, with the functional form preserved and only the coefficients changing. It also marks its boundary: cross-corpus prediction over-predicts code, implicating redundancy as a second axis of a broader $Φ(D,R,A,H)$ data-to-weight framework.

Comments27 pages, 14 figures, 5 tables. Code and data: https://github.com/tiexinding/NPM-Weibull-public

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑