arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解并利用训练后阶段中的各向异性

Understanding and Exploiting Anisotropy in Post-Training

Samyak Jha, Harshvardhan Saini, Yizhen Liao, Yiming Tang, Dianbo Liu

arXiv 2609.32792首次发表:更新:

AI 中文总结

本文揭示LLM训练后阶段中残差通道各向异性并非缺陷,而是连贯性基底与推理适应的分工,并提出SphereGate方法利用该特性,以极少参数超越参数高效基线并媲美全模型GRPO。

AI 中文摘要

LLM训练后阶段结合了监督微调(SFT)(一种覆盖模式的前向KL目标)与强化学习(RL)(一种寻找模式的逆向KL目标)。频率加权似然训练留下了一个众所周知的特征:各向异性,即少数残差通道承载了不成比例的大激活值。各向异性已被广泛记录,通常被视为缺陷,但其功能及其与训练后阶段的交互仍不清楚。我们首先对其进行分析。一个无标签的异常值规则隔离了约5%的残差通道,这些通道对语言建模至关重要:移除它们会使困惑度从10升至超过10^6,而移除数量匹配的随机通道仅使困惑度升至35。然而,这些通道几乎无法区分正确与错误的推理。SFT重塑这些通道,而RL则基本保持它们不变,并适应互补通道。因此,这些通道构成了模型的连贯性基底,而推理适应发生在其他地方。随后我们利用这一特性。SphereGate在冻结的主干网络上为每个残差通道学习一个有界增益。其激活加权梯度可证明地限制了高能量连贯性通道的移动,并让其余通道保持自由。仅用0.1M可训练参数,SphereGate在MATH-500上,在Qwen2.5(0.5B至7B)和Llama-3-8B上,优于参数高效基线2.0至7.3个百分点,与全模型GRPO相当或更优。各向异性不是缺陷,而是训练后阶段可利用的一种分工。

英文摘要

LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in which a few residual channels carry disproportionately large activations. Anisotropy is widely documented and usually treated as a defect, yet its function and its interaction with post-training remain unclear. We first analyze it. A label-free outlier rule isolates about 5\% of residual channels that are essential for language modeling: removing them raises perplexity from 10 to over $10^6$, versus 35 for count-matched random channels. Yet they barely distinguish correct from incorrect reasoning. SFT reshapes them, whereas RL leaves them largely intact and adapts the complementary channels. These channels therefore form the model's \emph{coherence substrate}, and reasoning adaptation happens elsewhere. We then exploit this. \textsc{SphereGate} learns one bounded gain per residual channel on a frozen backbone. Its activation-weighted gradients provably limit movement of high-energy coherence channels and leave the remaining channels free. With 0.1M trainable parameters, \textsc{SphereGate} outperforms parameter-efficient baselines by 2.0--7.3 points on MATH-500 across Qwen2.5 (0.5B--7B) and Llama-3-8B, is comparable or exceeds full-model GRPO. Anisotropy is not a defect but a division of labor that post-training can exploit.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑