arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个QK通道,多个源:防范低精度注意力崩溃

One QK Channel, Many Sources: Tracing Low-Precision Attention Collapse

Shuxiao Xie, Shuyang Xie, Yuan Cao, Dezhi Ran, Wei Yang, Tao Xie

arXiv 2608.02091首次发表:更新:

发表机构

Beijing Tongming Lake Information Technology Application Innovation Center (TLAIC); Fudan University Institute of Systems for Advanced Computing; Harbin Institute of Technology; Peking University; Shanghai Institute of Systems for Open Computing(北京通明湖信息技术应用创新中心(TLAIC); 复旦大学先进计算系统研究所; 哈尔滨工业大学; 北京大学; 上海开放计算系统研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对bfloat16变换器的低精度注意力崩溃问题,通过定位故障源与通道的分离关系,提出QK-Guard方法,可有效抑制注意力失控,支持在共享QK位点进行干预。

AI 中文摘要

bfloat16变换器可正常训练多步后突然崩溃,不同的低精度错误可触发相同故障,目前尚不清楚是否需为每个源单独修复,还是可阻断一条共享路径。我们将复现的GPT-2级崩溃定位到流式softmax累加器,其中fp32累加可修复该崩溃,并将此故障作为分析工具,用于跨源移动受控错误。置于注意力外部的错误仍会引发相同的查询-键(QK)谱失控,而仅校正QK可在源故障活跃时保持训练稳定。这种源-通道分离表明,故障源并非故障通道,该结论在测试的所有架构和规模中均成立,且在第二种GPU架构上也可复现。因果探针将每次更新从当前QK权重的前三个奇异方向上投影:查询投影的最大奇异值保持在11.1,而移除其他地方的同等能量后,该值升至237。因此,QK通道驱动早期失控,而非仅跟踪失控。进入该通道取决于各步骤间的时间符号一致性,而非总偏差。QK-Guard通过一个休眠控制器关闭该通道,当注意力logit饱和时,该控制器会开启无参数的QK归一化。它可抑制所有测试的失控,且在60k步以上的表现与始终开启的QK归一化相当,而在同一触发点采取的非QK操作则失败。研究结果支持在共享QK位点进行干预,而非对每个故障源单独修复。

英文摘要

A bfloat16 transformer can train normally, then collapse abruptly. Prior work links collapse to structured attention errors and shows QK normalization disrupts their compounding. Distinct low-precision errors trigger the same collapse, leaving unclear whether each needs a fix at its source or one shared route can be blocked instead. We isolated the fault behind a reproduced GPT-2-class collapse to the streaming-softmax accumulator, where an fp32 streaming core repairs it, and turned it into an assay for moving a controlled error across sources. Using it, we found that errors placed outside attention still drove the same QK spectral runaway, and that correcting only QK kept training stable while the fault stayed active. This is a source-channel dissociation: fault source is not failure channel. It held across tested architectures and scales, and reproduced on a second GPU architecture. As a causal probe, projecting each update off the current QK weights' leading three singular directions held the query projection's largest singular value to 11.1, whereas removing equal energy elsewhere left it at 237: the QK channel causally drives the early runaway. What lets the injected error in is temporal sign-coherence, its per-head sign persisting across steps, not aggregate deviation; once inside, the runaway shows as attention-logit saturation. QK-Guard, a dormant controller, tests this by switching on parameter-free QK normalization at the first monitored threshold crossing. On the runs designated for this test, the QK-local action prevented the failure of each matched or same-configuration unguarded run; on plain GPT-2, all 12 final train and validation losses were within 0.03 nat of same-configuration always-on QK-norm, and both methods ran 60k steps without collapse. Intervention at the QK locus therefore suffices in place of a fix at each source.

Comments24 pages, 4 figures. Revised title and presentation; corrected the streaming-core description and updated control analyses. Updated research artifact: https://github.com/xieTwim/one-qk-channel-artifact/releases/tag/v0.2.0-arxiv-v2

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑