arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当注意力失效:ALiBi位置编码中的数值故障

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

Christopher Schröder, Lukas Gienapp, Ferdinand Schlatt, Martin Potthast, Gerhard Heyer

arXiv 2608.03994首次发表:更新:

AI 中文总结

该研究发现ALiBi位置编码存在线性偏置缩放导致浮点下溢的失效模式,提出四种训练缓解策略,发现对数缩放距离在密码检索中改进最稳定,为ALiBi模型训练提供具体建议。

AI 中文摘要

我们发现了ALiBi位置编码此前被忽视的一种失效模式:其线性偏置缩放会导致浮点精度下溢,使大量注意力权重归零,导致受影响的注意力头部分“失明”。我们分析了该失效模式,明确了其影响,并研究了四种缓解策略。我们进一步证明,该问题在基于ALiBi的最先进预训练模型中确实存在。对1.48亿参数解码器模型进行的全面预训练实验,帮助我们将其影响与上下文外退化分离开来。我们发现,ALiBi的这种失效模式会严重损害令牌检索,而对标准解码器基准的影响较小。我们提出了四种训练时缓解策略,并对其单独及组合效果进行评估,发现对数缩放距离在密码检索中能带来最稳定的改进。尽管存在该问题,默认的ALiBi斜率仍是一个极强的基线,尤其在干草堆中的针检索任务上表现突出。基于这些发现,我们提供了使用ALiBi训练模型的具体建议。

英文摘要

We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes out a large fraction of attention weights and renders the affected attention heads partially blind. We analyze this failure mode, characterize its impact, and examine four mitigation strategies. We further demonstrate its occurrence in state-of-the-art pretrained models based on ALiBi. Comprehensive pretraining experiments with 148M-parameter decoder models help us to disentangle its effects from out-of-context degradation. We find that ALiBi's failure mode can substantially impair token retrieval while having only a minor effect on standard decoder benchmarks. We propose four training-time mitigation strategies and evaluate them individually and in combinations, finding that log-scaled distances yield the most consistent improvements in passkey retrieval. Despite this problem, default ALiBi slopes remain a surprisingly strong baseline, particularly for needle-in-a-haystack retrieval. Based on these findings we provide concrete recommendations on how to train models with ALiBi.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑