发表机构
INTI International College Penang; FPT University(槟城英迪国际学院; FPT大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究分析六种输入扰动在仅解码器语言模型中于三个层面的传播,发现单一鲁棒性评估度量存在误导性,提出需对大型语言模型鲁棒性开展多层面评估。
AI 中文摘要
语言模型会遇到拼写错误、损坏文本、修改后的词语以及被打乱的词元顺序,但其鲁棒性通常仅通过输出行为进行评估。我们研究六种自然及合成输入扰动如何在仅解码器架构的语言模型中传播,涉及三个层面:输出行为、隐状态几何结构以及注意力头功能。我们通过中心化核对齐和本征维度分析层状几何结构,评估四个GPT-2和两个Qwen2.5检查点的行为效应,并研究GPT-2中的注意力头响应。扰动类型产生可区分的度量特征,这些特征未被输出度量完全捕获,且在测试的检查点间仅部分一致。复制分数与词元替换和打乱下的激活修补恢复尤其相关。在GPT-2中,梯度引导的HotFlip扰动比速率匹配的随机词元替换引发更强的行为和表征破坏;其行为效应在所有六个测试检查点间一致。我们的结果表明,基于单一行为或表征度量的鲁棒性主张可能具有误导性,推动对扰动如何改变语言模型计算进行多层面评估。
英文摘要
Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at three levels: output behavior, hidden-state geometry, and attention-head function. We evaluate behavioral effects across four GPT-2 and two Qwen2.5 checkpoints, analyze layerwise geometry using centered kernel alignment and intrinsic dimension, and examine attention-head responses in GPT-2. Perturbation types produce distinguishable metric profiles that are not fully captured by output measures and are only partly consistent across the tested checkpoints. Copying scores show the strongest pooled associations with activation-patching recovery under token substitution and shuffling, although these associations do not isolate copying-specific effects. Gradient-guided HotFlip perturbations also cause stronger behavioral and representational disruption than rate-matched random token substitutions in GPT-2; their behavioral effects are consistent across all six tested checkpoints. Our results show that robustness claims based on a single behavioral or representational metric can be misleading, and motivate multi-level evaluation of how perturbations alter language-model computation.
Comments15 pages, 6 figures; Accepted at the NeurIPS 2026 InterpScience workshop