arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03842cs.CLcs.LG

敏感性、因果性与修复能力的分离:扰动鲁棒性的逐层分析及其缩放规律

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

Nathan Labiosa, David Buff, Ena Nayak, Erica Donno

AI总结:

该研究分析了语言模型的扰动鲁棒性,发现敏感性、因果性、补偿能力三种层映射分离,识别两种传播机制,提出级联中断机制,给出实用指导并指出方法学警示。

AI中文摘要:

当语言模型在经过表面扰动的输入(拼写错误、OCR噪声、同音异义词)上失效时,“哪一层负责”有三种自然的可操作定义:表示差异最大的层(敏感性)、恢复干净激活可恢复预测的层(因果性)、小型适配器可修复损伤的层(补偿能力)——我们证明这三种层映射是分离的。在包含五个模型的面板中,我们识别出两种传播机制:尖峰抑制型(Phi-3.5、Gemma-2-9B)和后期累积型(Llama-3、Mistral、Qwen2.5-7B);在满足80%身份补丁阈值的两个模型上,敏感性与因果性呈负相关(相关系数ρ=-0.72至-0.88)。对Qwen2.5(1.5B至14B)的家族内缩放显示,后期累积特征随模型规模单调增强,且在第二个模型家族中得到验证。我们提出级联中断作为分离现象的机制:放置在因果相关早期层的适配器会破坏完整的下游计算,使得诊断标记的位置成为最差的适配器放置点。对四个模型(3.8-8B)的固定框架层扫描在思维链GSM8K任务上证实了核心预测——标记位置是每个可判定模型上最具破坏性的适配器窗口;在多项选择对照任务上则符号一致但强度大幅减弱,这与损伤随生成长度累积的结论一致。该扫描给出了实用指导:无需训练的LRD预筛选和默认最深层放置规则,尽管与无适配器基线相比的绝对增益仍然很小。最后,在充足的生成预算下,表示稳定性损失带来的表观增益会反转——截断的思维链曾被计为空值——这对任何在思维链任务上评估的干预措施都是一个方法学警示。

英文摘要:

When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate. Across a five-model panel we identify two propagation regimes - spike-and-suppress (Phi-3.5, Gemma-2-9B) and late-accumulation (Llama-3, Mistral, Qwen2.5-7B) - and on the two models meeting an 80% identity-patch gate, sensitivity and causality are anti-correlated (rho = -0.72 to -0.88). Within-family scaling on Qwen2.5 (1.5B to 14B) shows the late-accumulation signature strengthening monotonically with scale, corroborated on a second family. We propose cascade disruption as the mechanism behind the dissociation: adapters placed at causally implicated early layers break intact downstream computation, making diagnostic-flagged sites the worst adapter placements. A fixed-harness layer sweep across four models (3.8-8B) confirms the core prediction on chain-of-thought GSM8K - the flagged sites are the most damaging adapter windows on every adjudicable model - and is sign-consistent but strongly attenuated on a multiple-choice control, consistent with damage that compounds with generation length. The sweep yields practical guidance: a training-free LRD pre-screen and a default-deepest placement rule, though absolute gains over no-adapter baselines remain small. Finally, apparent gains from a representation-stability loss reverse under an adequate generation budget - truncated chain-of-thought had been scored as empty - a methodological warning for any intervention evaluated on chain-of-thought tasks.

补充信息

↑