贝叶斯重复惩罚:一种用于扭转自回归语言模型中注意力崩溃的有原则的相邻条件框架
Bayesian Repetition Penalty: A Principled Adjacent-Conditional Framework for Reversing Attention Collapse in Autoregressive Language Models
浏览论文内容
中文总结 AI 辅助
针对自回归语言模型注意力崩溃问题,提出基于相邻条件概率构造的贝叶斯重复惩罚框架,通过自归一化惩罚率校正异常置信度,经指数移动平均累积到输出层偏差,实验表明该机制能拯救崩溃模型并保持生成质量。
中文摘要 AI 辅助
自回归语言模型中的注意力崩溃表现为重复令牌循环,即模型陷入自我强化吸引子,这是现有解码时启发式方法无法从根本上解决的持续问题。我们提出了一个有原则的框架,通过相邻条件概率构造将令牌的观察频率与其语料库先验进行比较,对崩溃生成模式产生的异常置信度进行惩罚或补偿。由此产生的自归一化惩罚率$R=f(m,n,p)/f(np,n,p)$无需特殊标准化,且具有零近似误差的封闭形式对数偏移。校正与损失梯度隔离,并通过指数移动平均累积到冻结的输出层偏差中,可作为已崩溃模型的修复机制,无需对标准训练管道进行侵入性修改。在一个15亿参数模型上的实验验证表明,冻结偏差机制可以拯救已陷入崩溃吸引子的模型,将2-gram重复率从0.073降至接近0,同时保持生成质量。
英文摘要
Attention collapse in autoregressive language models -- manifested as repetitive token loops where the model becomes trapped in self-reinforcing attractors -- is a persistent pathology that existing decoding-time heuristics fail to address at its root cause. We present a principled framework that penalises or compensates anomalous confidence arising from collapsed generation patterns, by comparing a token's observed frequency against its corpus prior through an adjacent-conditional probability construction. The resulting self-normalising penalty ratio $R=f(m,n,p)/f(np,n,p)$ requires no ad hoc standardisation and admits a closed-form logit offset with zero approximation error. The correction is isolated from the loss gradient and accumulated into a frozen output-layer bias via exponential moving average, enabling deployment as a repair mechanism for models that have already collapsed without requiring intrusive modifications to standard training pipelines. Experimental validation on a 1.5B-parameter model demonstrates that the frozen-bias mechanism can rescue a model already trapped in a collapsed attractor, reducing 2-gram repetition from 0.073 to near 0 while preserving generation quality.