arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

注意力敏感性并非足够:微调下的注意力层级与行为式上下文学习的分离

Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning

Jinyuan Zhang, Peng He, He Hu, Yin Yuan, ShengShuo Jiao

arXiv 2609.00064首次发表:更新:

发表机构

Hubei University(湖北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对Llama-2-7B,发现微调下注意力层级的ICL代理与实际行为分离,提出ICS指标,验证其需经行为差距验证才适合作为训练目标。

AI 中文摘要

上下文学习(In-Context Learning, ICL)让大语言模型可通过示例适配新任务,而微调会削弱该行为。许多保留诊断方法会检查注意力:若示例改变时注意力发生变化,模型被视为上下文敏感。本文探究该代理在优化后可信任的程度。我们形式化了上下文敏感性(In-Context Sensitivity, ICS),即匹配与不匹配示例前缀上最后一个标记注意力的平均行距离,并将其与ICL-GAP配对,ICL-GAP是相同前缀间的行为准确率差距。在对Llama-2-7B的受控四组 ablation 实验中,最大化ICS的正则化项($\text{armKL}$)将ICS提升至1.413,距其几何上限仅0.5%。但行为读数给出不同结果:ICL-GAP保持接近零,MMLU准确率从0.371降至0.279,这是有界注意力代理的古德哈特定律式分离。端点统计定位了机制:注意力在各前缀间变得尖锐且近乎不相交,但路由至格式和示例主体标记而非标签。随机标签协议确认,行为探测系列在相同检查点仍保留动态范围。在构造性扫描中,行为门控部分缓解了该效应,而锚定预训练计算的目标则保持了高MMLU、中等ICS的区域,这是发散最大化者留下的区域。主要的诊断教训是:注意力层级的ICL代理仅在经行为差距验证后,才有资格作为训练目标。

英文摘要

In-context learning (ICL) lets large language models adapt to new tasks from demonstrations, and fine-tuning can erode this behaviour. Many preservation diagnostics inspect attention: if attention changes when demonstrations change, the model is treated as context-sensitive. This paper asks how far that proxy can be trusted once it is optimised. We formalise \emph{In-Context Sensitivity} (ICS), the average row distance between last-token attention on matched and mismatched demonstration prefixes, and pair it with \emph{ICL-GAP}, the behavioural accuracy gap between the same prefixes. In a controlled four-arm ablation on Llama-2-7B, an ICS-maximising regulariser ($\armKL$) drives ICS to $1.413$, within $0.5\%$ of its geometric ceiling. The behavioural readout tells a different story: ICL-GAP stays near zero and MMLU accuracy moves from $0.371$ to $0.279$, a Goodhart dissociation of the bounded attention proxy. Endpoint statistics locate the mechanism: attention grows sharp and near-disjoint across prefixes yet routes to formatting and demonstration-body tokens rather than labels. A random-label protocol confirms that the behavioural probe family retains dynamic range at the same checkpoints. In a constructive sweep, behaviour gating partially mitigates the effect, while objectives anchored to pretrained computation hold the high-MMLU, moderate-ICS region that divergence maximisers leave. The main lesson is diagnostic: attention-level ICL proxies earn their place as training targets only after validation against behavioural gaps.

Comments15 pages, 5 figures; appendices included

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑