arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35475cs.CLcs.AI

自发上下文恢复:语言模型如何从受损输入中恢复

Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs

  • Independent Researcher(独立研究者)
  • ML Collective (MLC)(机器学习集体(MLC))

机构由 AI 辅助整理,请以论文原文为准。

Pranjal Garg, Jacob Beck

AI总结:

本文研究语言模型在输入损坏时自发恢复正确输出的内部机制,发现两阶段过程,并利用隐藏状态预测失败,通过微调提升鲁棒性。

AI中文摘要:

语言模型有时即使输入被删除、替换或拼写错误所破坏,也能产生正确的输出。我们研究了伴随这种行为的内部过程,称之为上下文恢复,在受控的仅注意力变换器和五个预训练的大语言模型(1B-32B参数)上,涵盖算术、阅读理解和多项选择推理任务。在仅注意力变换器中,尽管仅在干净序列上训练,没有损坏训练或显式去噪目标,恢复仍然自发出现。我们发现上下文恢复遵循两阶段过程:早期层在受损位置定位与修复相关的效应,而后期层通过残差流在未受损位置累积这些效应,并最终在输出位置集中它们。修复结果可以从隐藏状态预测:在仅注意力模型中,与干净状态的余弦对齐具有高度预测性,而线性探针在预训练的大语言模型中恢复了额外信息。仅使用受损提示的第一块隐藏状态的线性探针预测失败的平均ROC-AUC为0.78。这使得在匹配或部分偏移的部署条件下进行失败分诊成为可能,并可能减少不必要的验证或计算。失败的示例还显示出沿损坏方向的显著更大的非线性。中等损坏的微调增加了损坏容忍度,同时减少了位移归一化线性化误差,将改进的鲁棒性与对损坏更接近线性的响应联系起来。

英文摘要:

Language models sometimes produce correct outputs even when their inputs are corrupted by deletion, replacement, or misspelling. We study the internal processes accompanying this behavior, which we call context restoration, in controlled attention-only transformers and five pretrained LLMs (1B-32B parameters) across arithmetic, reading comprehension, and multiple-choice reasoning tasks. In the attention-only transformers, restoration emerges spontaneously despite training exclusively on clean sequences, without corruption training or an explicit denoising objective. We find that context restoration follows a two-phase process: early layers localize effects associated with repair at corrupted positions, while later layers accumulate these effects at uncorrupted positions through the residual stream and ultimately concentrate them at the output position. Repair outcome is predictable from hidden states: cosine alignment with the clean state is highly predictive in attention-only models, while linear probes recover additional information in pretrained LLMs. A linear probe using only the corrupted prompt's first-block hidden state predicts failure with mean ROC-AUC 0.78. This enables failure triage under matched or even partially shifted deployment conditions and may reduce unnecessary verification or computation. Failed examples also show substantially greater nonlinearity along corruption directions. Moderate-corruption finetuning increases corruption tolerance while simultaneously reducing displacement-normalized linearization error, associating improved robustness with a more nearly linear response to corruption.

↑