AI 中文总结
研究检索任务中注意力与答案因果依赖的关系,发现二者常不一致。提出用因果证据集监督稀疏注意力路由,无需注释,能让选择器训练到更高准确率,在预训练模型中也验证了该方法的有效性。
AI 中文摘要
稀疏注意力通过让每个查询仅读取输入的选定部分来降低长上下文的成本。这些选择器通常通过提炼密集教师的注意力模式来训练,假设注意力揭示了教师实际使用的上下文。我们在每个答案的证据已知的检索任务上测试了该假设。通过掩盖部分上下文并测量答案是否改变,我们发现注意力和因果依赖常常不一致,并且提炼的选择器继承了这种不匹配。教师关注他们已学会忽略的过时事实,并且即使他们依赖相同的证据,他们的注意力在不同训练运行中也可能变化。在两步参考任务中,答案处的注意力跳过中间步骤,因为它在前向传播中更早解决:基于注意力训练的选择器准确率为41%,而基于因果证据训练的相同选择器达到99%并与教师匹配。这些证据集无需注释:仅通过掩盖从冻结的教师中恢复,它们将选择器训练到相同的准确率。我们在预训练模型中也发现了同样的冲突:Qwen2.5 - 3B在58%的冲突事实示例中对过时事实的关注多于当前事实,尽管回答正确,而Gemma - 2 - 9B在限制为两个相关句子时准确率从56%提高到99%。注意力显示了模型查看的位置,不一定是其答案所依赖的内容;在我们测试的各种情况下,这种依赖作为训练目标与注意力匹配或优于注意力。
英文摘要
Pruning a long context means committing to the blocks a model will keep, and the usual selector is distilled from a dense teacher's attention. That assumes attention shows which context the answer depends on. We test the assumption on retrieval tasks where the evidence is known exactly, by masking context and measuring whether the answer changes. Attention and causal dependence disagree. Teachers attend to outdated facts that the answer does not depend on, and they attend differently across training runs that use the same evidence. Selectors trained on that attention copy both failures. On a multi-hop retrieval task, a selector distilled from attention routes at 36% to 98% depending on the training run. The same selector trained on causal evidence sets reaches 99% or better on every run. Dense accuracy does not tell the teachers apart. Masking the frozen teacher recovers the causal sets of these tasks without annotations. Frozen pretrained models show the same conflict, and selectors supervised with known evidence labels beat attention-based eviction through 32B when context must be pruned before the question arrives.
Comments36 pages, 6 figures. v4: adds a pretrained-scale block-eviction appendix: causal- vs attention-supervised routers on RULER at 8K-32K across Qwen 3B-32B and Llama-3.1-8B with controls and seed replications, a QASPER natural-text study with intervention-derived labels, and a reconstruction eviction baseline. Main text updated; code in ancillary files