arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

词汇扰动破坏大语言模型推理:注意力转移的实证研究

Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

Jiaqian Zhu, Yang Zhang, Junhua Ding, Xiaowei Yu

arXiv 2608.22140首次发表:更新:

发表机构

Missouri University of Science and Technology; University of North Texas(密苏里科技大学; 北得克萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过实证分析发现词汇扰动会破坏LLM的推理能力,揭示了注意力转移机制及词元内容与注意力分配的耦合关系,解释了现有推理策略难以修复性能的原因。

AI 中文摘要

大语言模型(LLMs)具备强大的推理能力,但其对现实中词汇损坏的鲁棒性仍未得到充分理解。我们在键盘噪声、字符交换和填充词插入三种场景下,针对四个推理基准,评估了四个开放权重指令微调模型和前沿模型。字符级扰动会大幅降低准确率,尤其是在多步推理任务中,而填充词插入的影响很小。我们将这种不对称性归因于注意力转移:词汇损坏会碎片化子词分词,产生的碎片会吸引不成比例的注意力权重,且集中在Transformer的中间和最终层。长度匹配的对照实验证实,导致性能损失的是碎片化而非提示长度。随后的因子干预实验揭示了损坏难以修复的原因:碎片化同时损坏了词元内容和注意力分配,且二者相互耦合。在内容损坏时恢复干净的注意力会产生负面影响,仅恢复内容则不够,只有同时恢复二者才能恢复大部分性能差距。这种耦合解释了为何推理时策略(包括思维链提示、拼写检查、自我修复及更强的修复模型)无法持续恢复性能:每种策略仅处理了其中一个通道。代码和数据可在该https URL获取。

英文摘要

Large Language Models (LLMs) achieve strong reasoning performance, but their robustness to realistic lexical corruption remains poorly understood. We evaluate four open-weight instruction-tuned models and frontier models across four reasoning benchmarks under keyboard noise, character swaps, and filler insertion. Character-level perturbations substantially degrade accuracy, especially on multi-step reasoning tasks, while filler insertion has little effect. We trace this asymmetry to Attention Diversion: lexical corruption fragments subword tokenization, and the resulting fragments attract disproportionate attention mass, concentrated in middle and final transformer layers. Length-matched controls confirm that fragmentation, not prompt length, drives the loss. A factorial intervention then shows why the damage is hard to undo: fragmentation corrupts token content and attention allocation together, and the two are coupled. Restoring clean attention while the content remains corrupted is actively harmful, restoring content alone is insufficient, and only restoring both recovers a substantial share of the gap. This coupling explains why inference-time strategies, including chain-of-thought prompting, spell-checking, self-repair, and stronger repair models, fail to consistently recover performance: each addresses one channel at a time. Code and data are available at https://github.com/Jiaqian-Janelle/Attention-Diversion

CommentsAccepted to EMNLP 2026 (Main Conference). 9 pages main text, 12 figures, 20 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑