发表机构
Kioxia Corporation(铠侠公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过反例证明,在残差补全中,局部重建增益的提升并不必然带来最终模型输出的保真度改善,精确恢复反而更有效。
AI 中文摘要
残差补全通过估计从精确稀疏计算中省略的令牌的贡献来增强查询感知的稀疏注意力。我们探究在相同的输入Q/K/V和所选支持集上,改进某一层的注意力输出重建是否必然能提高最终模型输出的保真度。我们研究了无需训练的RESA以及带有冻结骨干语言模型的学习型Top-K+φ方法。一个预设的单层筛选产生了两个Qwen3-0.6B/Multi-LexSum干预案例,在这些案例中,直接运行时测量显示预设的请求聚合局部重建增益为正,但在发现集和提示令牌不相交的保留请求上,最终KL保真度均差于对应的全弃权(不执行)精确Top-K基线。在同一层进行精确恢复反而改善了最终保真度,表明在这些案例中,这种逆转是近似补全所特有的。在补充的多层实验中,一个与任务无关的局部诊断通常能修复所测试的补全估计器,尽管修复后的模型并未持续优于精确Top-K。综合这些结果,表明更好的局部重建并不必然转化为更好的最终模型保真度。
英文摘要
Residual completion augments query-aware sparse attention by estimating the contribution of tokens omitted from the exact sparse computation. We ask whether improving a layer's attention-output reconstruction on the same incoming Q/K/V and selected support necessarily improves the fidelity of the final model output. We study training-free RESA and learned Top-K+$ϕ$ with frozen backbone language models. A prespecified single-layer screen yields two Qwen3-0.6B/Multi-LexSum interventions for which direct-runtime measurements show positive prespecified request-aggregate local reconstruction gain but worse final KL fidelity than the corresponding all-abstain Exact Top-K baseline on both discovery and prompt-token-disjoint holdout requests. Exact restoration at the same layer instead improves final fidelity, showing that the reversal is specific to approximate completion in these cases. In complementary multi-layer experiments, a task-independent local diagnostic often repairs the tested completion estimators, although the repaired models do not consistently outperform Exact Top-K. Together, these results show that better local reconstruction need not translate into better final-model fidelity.