arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当排名随LLM退化而上升

When Rank Rises as LLMs Degrade

Zhaohui Geoffrey Wang

arXiv 2610.09647首次发表:更新:

发表机构

University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示LLM后训练中谱统计量方向随机制变化,RankMe等指标可能误判退化,提出双侧多通道监测器可诊断损伤但无法提前预警。

AI 中文摘要

后训练在非平稳环境中适应语言模型。从业者使用RankMe及相关谱统计量监测表征健康状况,通常假设当表征退化时排名会下降。我们表明,这一假设对于LLM后训练是不安全的。在一项针对Qwen3-0.6B的受控研究中,包含四种退化模式和三种随机种子,数据重复使留出损失相对于健康状态恶化75%,同时增加了原始和中心化的RankMe;后者变化了13.5个合并标准差。协方差有效秩上升到接近其健康值的两倍。这种失败是谱离散而非坍缩,因此单侧监测器将最差的检查点评为最健康的。相比之下,学习率配置错误降低了中心化RankMe和k95,而未中心化RankMe在不同种子间不一致。因此,方向是机制-统计量对的属性,不能仅通过重新校准来修复。我们还区分了两个常被混淆的统计量:RankMe对奇异值进行归一化,而协方差有效秩对特征值进行归一化。在预训练模型的原始中间层状态上,大规模激活将后者固定在维度d附近的1,而RankMe保留了可用范围。然后,我们测试了一个双侧、多通道顺序监测器,使用独立的校准和测试数据。在预注册的共享前缀、留一种子评估中,它在分叉后10到60步的每个折中检测到所有三种损伤机制,并通过触发方向区分离散和向下排名损伤。然而,它从未早于留出探针损失,且使用两种子校准会在留出的健康种子上产生误报。谱监测可以诊断失败机制,但它不会比留出损失更早警告,且有效性声明需要留出的健康数据。

英文摘要

Post-training adapts language models in non-stationary environments. Practitioners monitor representation health with RankMe and related spectral statistics, often assuming that rank falls when representations degrade. We show that this assumption is unsafe for LLM post-training. In a controlled study of Qwen3-0.6B with four degradation modes and three seeds, data duplication worsens held-out loss by 75% relative to healthy while increasing both original and centred RankMe; the latter changes by 13.5 pooled standard deviations. Covariance effective rank rises to nearly twice its healthy value. This failure is spectral dispersion rather than collapse, so a one-sided monitor rates the worst checkpoint as the healthiest. By contrast, a learning-rate misconfiguration lowers centred RankMe and k95, while uncentred RankMe is inconsistent across seeds. Direction is therefore a property of the regime-statistic pair and cannot be fixed by recalibration alone. We also distinguish two often-conflated statistics: RankMe normalises singular values, whereas covariance effective rank normalises eigenvalues. On raw intermediate-layer states in the pretrained model, massive activations pin the latter near 1 out of dimension d while RankMe retains usable range. We then test a two-sided, multichannel sequential monitor with separate calibration and test data. In a pre-registered shared-prefix, leave-one-seed-out evaluation, it detects all three damage regimes in every fold 10 to 60 steps after the fork and separates dispersion from downward-rank damage by firing direction. However, it never precedes held-out probe loss, and calibration with two seeds produces false alarms on the held-out healthy seed. Spectral monitoring can diagnose failure regimes, but it does not warn earlier than held-out loss, and validity claims require held-out healthy data.

CommentsNeurIPS 2026 Workshop on Continual Learning for Foundation Models and Agents (CL4FMAgents); 8 pages + appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑