arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对齐悖论:后训练如何放大语言模型中的高置信幻觉

The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models

Qingjia Huang, Yakai Li, Jianguo Wu, Qihang Zhou, Aimin Yu, Xiaoqi Jia, Luping Ma, Weijuan Zhang

arXiv 2609.32617首次发表:更新:

AI 中文总结

研究发现后训练对齐本身是放大语言模型高置信幻觉的主因,提出在DPO中引入熵依赖边际界,将Mistral-7B的高置信错误减少达35.3%。

AI 中文摘要

大型语言模型(LLMs)可能以高置信度产生事实上不正确的答案,这削弱了其可靠性,并限制了基于不确定性的错误检测的有效性。虽然先前的研究将高置信幻觉归因于训练数据中知识缺失、推理错误或随机解码等因素,但我们发现后训练对齐本身是这些错误的主要驱动因素,我们将这一现象称为“对齐悖论”。在五个模型家族的事实性基准评估中,未对齐的基础模型在长尾事实查询上产生的高置信错误很少,而指令微调模型将高置信错误(p ≥ 0.95)放大了超过一个数量级(10倍至35倍)。使用Logit Lens进行逐层探测显示,这种过度自信出现在深层,其中错误答案的边际在早期和中间层保持接近零后,在深层扩展到超过4.0点。这些发现促使在后训练期间限制边际增长。我们通过直接偏好优化(DPO)中的熵依赖边际界来实现这一原则。在使用Mistral-7B的多轮实验中,有界目标相对于标准DPO将高置信错误减少了高达35.3%,同时保持了在评估的一般推理基准上的性能。这些结果表明,有界边际在后训练期间缓解了高置信幻觉。

英文摘要

Large language models (LLMs) can produce factually incorrect answers with high confidence, undermining their reliability and limiting the effectiveness of uncertainty-based error detection. While prior research attributes confident hallucinations to factors such as missing knowledge in training data, reasoning errors, or stochastic decoding, we uncover that post-training alignment itself is a primary driver of these errors, a phenomenon we call the \textbf{Alignment Paradox}. Across five model families evaluated on factual benchmarks, unaligned base models produce few high-confidence errors on long-tail factual queries, whereas instruction-tuned models multiply high-confidence errors ($p \ge 0.95$) by more than an order of magnitude (10$\times$ to 35$\times$). Layer-wise probing with the Logit Lens reveals that this overconfidence emerges in late layers, where wrong-answer margins expand past 4.0 points after remaining near zero across early and intermediate layers. These findings motivate limiting margin growth during post-training. We implement this principle through an entropy-dependent margin bound in direct preference optimization (DPO). In multi-epoch experiments with Mistral-7B, the bounded objective reduces high-confidence errors by up to 35.3\% relative to standard DPO while maintaining performance on evaluated general reasoning benchmarks. These results show that bounded margins mitigate confident hallucinations during post-training.

CommentsCode: https://github.com/star5o/HCE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑