arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可靠性呈反比:更大的模型通过隐藏的自回归风险机制更快地加剧错误

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models

Kushal Chakrabarti

arXiv 2607.18292首次发表:更新:

发表机构

Obviously Wrong, LLC(明显错误有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现随着语言模型规模扩大,答案虽更接近真实但退化更快。通过逐位置分歧追踪自回归风险残余,揭示了知识差距下降、知识退化增长等四个发现,指出更大模型会通过特定风险机制更快加剧错误,该机制因果且模型自身难以察觉。

AI 中文摘要

随着语言模型规模的扩大,答案开始更接近真实,但退化速度更快:扩大规模带来了能力,但侵蚀了可靠性。知识差距理论(更多数据、检索或规模)忽略了一种自回归风险残余,而规模会加剧这种残余:模型采用了低概率令牌,以此为基础并像滚雪球一样越滚越大。我们通过针对更强的同家族预言机的逐位置分歧δ=log p_M - log p_O来追踪这一情况,其二阶矩精确地分解为偏差²KL(p_M || p_O)²和风险Var[δ]。我们提出了四个发现:(i)在扩大规模时,知识差距下降了约6倍,而知识退化增长了11 - 39倍;(ii)在一次生成中,感知到的不确定性H(p_M)迅速放松,而预言机参考风险持续长达17倍之久,留下了一个自信但不稳定的风险机制,连接连续的生成(140亿参数模型时增加69%);(iii)这种机制是因果性的——一种策略性的、固定KL方差收缩在三个模型家族中减少了35 - 74%的网络验证幻觉;(iv)它在结构上规避了自我监控,仅基于p_M的探测器(如语义熵)在风险分支上触发的次数减少了约30%(p<10^-16),而该分支上的生成几乎多了4倍。更大的模型通过一种占主导地位、自我延续、因果性且模型自身不可见的失败模式更快地使错误像滚雪球一样越滚越大。

英文摘要

Bigger language models are less reliable. Across three families, three benchmarks and six rungs, including in-the-wild chat logs, scaling closes the start-of-response knowledge gap up to $7\times$ while within-response knowledge degradation grows up to $39\times$. We trace that residual to one variable, the per-position disagreement $δ= \log p_M - \log p_O$ against a stronger oracle, whose second moment splits exactly into bias$^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ and decoding risk $\mathrm{Var}[δ]$. That split is an interpretability statement before it is a statistical one: the model's self-readable uncertainty $H(p_M)$ enters only the bias term, so the risk term has no model-readable component. Risk also takes a growing share of the squared error with scale, $31\%$ to $49\%$ from $1.7$B to $14$B. At a fabrication $H(p_M)$ relaxes within one token while risk persists up to $23\times$ longer, leaving a confident-but-precarious regime that bridges consecutive fabrications ($+69\%$ at $14$B). Contracting that risk at fixed $\mathrm{KL}$ removes $35$-$74\%$ of web-verified hallucinations across six rungs and three families. Semantic entropy fires $\approx$$30\%$ less on that branch ($p\!<\!10^{-16}$) though it carries nearly $4\times$ the fabrications. Bigger models snowball mistakes faster, through a failure mode that is dominant, self-perpetuating, causal and invisible to the model itself.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑