arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文投毒作为长上下文语言模型中的极值注意力干扰

Context Poisoning as Extreme-Value Attention Interference in Long-Context Language Models

Meysam Ghaffari, Nina Fatehi, Bhaskar Sen, Nasim Sabetpour, Carlos Morato

arXiv 2609.22101首次发表:更新:

AI 中文总结

本研究提出上下文投毒概念,将其建模为注意力极值干扰,证明证据边际需随干扰项数量平方根增长,并通过实验验证了检索退化机制及缓解方向。

AI 中文摘要

大型语言模型能够处理越来越长的提示,但随着无关或混淆上下文的加入,它们定位和使用关键证据的能力可能会下降。我们将这一现象称为上下文投毒,并将其表述为注意力中的极值干扰:关键证据的得分有上界,而有效干扰项中的最大得分随其数量增长。在softmax检索抽象下,我们推导出一个有限样本上界,表明要在基础比率之上维持固定的准确率目标,证据边际需要按$\Omega(\sqrt{\log N})$的规模增长,其中$N$表示有效干扰项数量,而不一定是原始上下文长度。该分析将长上下文退化与得分混叠、位置混叠和softmax稀释联系起来。受控实验表明,在嵌入硬负样本的情况下,随着总上下文增长,检索准确率下降;在固定上下文长度下,相同格式条件在测试的干扰项构造中产生最大的观测准确率下降;检索门控可以改善证据使用,但其净收益取决于保留证据召回。这些结果促使了证据瓶颈、抗混叠表示、先检索后推理架构、验证器介导的记忆以及对比抗投毒训练。

英文摘要

Large language models can process increasingly long prompts, yet their ability to locate and use decisive evidence may degrade as irrelevant or confusable context is added. We formulate this phenomenon, which we call context poisoning, as extreme-value interference in attention: the decisive-evidence score is upper-bounded, while the maximum score among effective distractors grows with their number. Under a softmax retrieval abstraction, we derive a finite-sample upper bound showing that maintaining a fixed accuracy target above base rate requires the evidence margin to scale as $Ω(\sqrt{\log N})$, where N denotes the effective distractor count rather than necessarily the raw context length. The analysis connects long-context degradation to score aliasing, positional aliasing, and softmax dilution. Controlled experiments show that retrieval accuracy decreases as total context grows in the presence of embedded hard negatives, that the same-format condition produces the largest observed accuracy drop among the tested distractor constructions at fixed context length, and that retrieval gating can improve evidence use while its net benefit depends on preserving evidence recall. These results motivate evidence bottlenecks, alias-resistant representations, retrieve-then-reason architectures, verifier-mediated memory, and contrastive anti-poison training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑