arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01676cs.CLcs.LG

通过反事实评估理解长上下文基础模型中的稀疏注意力选择性

Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation

Xingyu Ren, Youran Sun, Chugang Yi, Haizhao Yang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过反事实评估探究长上下文基础模型的稀疏注意力选择性,证实稀疏化会改变内容影响力,提出开放测量框架,相关结果无法通过聚合准确率检测。

中文摘要 AI 辅助

稀疏注意力被广泛应用于长上下文服务栈中,但尚无框架可评估丢弃注意力块如何改变特定内容对模型输出的影响。我们首先证实该现象真实且具有因果性:在四种架构上的块级稀疏Flash注意力(Block Sparse Flash Attention, BSFA)路径重放,使16个单元中的13个改变了输出决策,且未出现身份重放标签翻转。随后,我们引入一种利用匹配探测卡的密集校准反事实评估,探测卡包括携带正确答案标签的Gold卡、携带目标错误标签的Poison卡以及仅含填充内容的Benign卡,在6种布局位置对称性下分离出稀疏化特有的影响。存在两种相互竞争的模式:信号集中性——选择器保留Gold块和Poison块的程度远高于与填充匹配的Benign块(在所有模型-任务对中G≈P≫B);整合损失——丢弃块会切断跨块注意力,这一结论通过一项消融实验得到证实,该实验中隔离探测块使其影响力从4.48 logits降至零。压缩率决定两者的平衡:对四种模型-任务对从轻度(c=0.25)到激进(c=0.75)压缩的完整扫描显示,四分之三的单元在更高压缩率下向更强的稀疏放大方向移动,其中两个单元出现符号反转。三个独立分支——BSFA路径重放、受控块top-k以及KV缓存驱逐——结果一致:稀疏化会改变内容影响力,而这种变化无法通过聚合准确率检测到。我们提供了一个可部署在任何暴露块身份的模型上的开放测量框架。

英文摘要

Sparse attention is widely deployed in long-context serving stacks, yet no framework audits how discarding blocks changes the influence of specific content on model output. We first establish that the phenomenon is real and causal: Block Sparse Flash Attention (BSFA) route replay across four architectures changes output decisions in 13 of 16 cells, with zero identity-replay label flips. We then introduce a dense-calibrated counterfactual audit using matched probe cards---Gold (carrying the correct answer label), Poison (carrying a target wrong label), and Benign (filler only)---under six-layout position symmetry, isolating the sparsification-specific effect. Two patterns compete. Signal concentration: the selector preserves Gold and Poison blocks far above filler-matched Benign blocks (G$\approx$P$\gg$B across all model--task pairs). Integration loss: discarding blocks severs cross-block attention---confirmed by an ablation where isolating the probe block collapses its influence from 4.48 logits to zero. Compression ratio governs the balance: a full sweep from mild ($c=0.25$) to aggressive ($c=0.75$) compression across four model--task pairs reveals that three of four cells move toward stronger sparse amplification at higher compression, with two exhibiting sign reversals. Three independent arms---BSFA route replay, controlled block-top-$k$, and KV-cache eviction---converge: sparsification changes content influence in ways aggregate accuracy cannot detect. We provide an open measurement framework deployable on any model exposing block identities.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)
  • University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

↑