AI 中文总结
该研究针对错误惩罚强化学习中弃权动作导致的奖励梯度与KL锚点同时失效的问题,提出通过训练强制置信度报告的结构性修复方案,实验验证了机制并提升了语言模型的相关性能。
AI 中文摘要
错误惩罚评分规则(正确回答得+1,错误回答得-λ,弃权得0)被越来越多地用于抑制幻觉:面对该规则的理性智能体,仅当正确概率超过Chow阈值t*=λ/(1+λ)时才会作答。我们证明,基于KL锚点的梯度学习器可能反其道而行之。当弃权是离散动作时,奖励梯度与锚点的恢复力会被同一门限饱和因子抑制并同时消失:在明确条件下(包括盲目作答会损失期望分数、提示具有有限读出量),模型会倾向于拒绝所有内容,其平均训练奖励随训练时间t以1/t的速率趋近于0,因此曲线看似在提升,但覆盖率已崩溃。优势估计器加剧了这一问题:在其稀疏作答机制下,组归一化会悄然将所有设计的惩罚替换为有效惩罚1,使学习到的阈值从t*变为1/2。修复方案是结构性的:训练一个强制置信度报告,采用严格正则评分加正确性奖励,仅在部署时通过阈值化该报告来执行弃权。始终输出的报告没有可饱和的门限,因此没有共享因子能同时破坏其奖励梯度和锚点,其校准后的最优解具有吸引力。模拟结果验证了所有预测,且在两个规模的语言模型上的实验证实了该机制的存在:该规则会使模型在优化器十步内本可解决的问题被抑制, ablation实验分离出了该原因,报告级训练同时提升了覆盖率、准确率和校准度。
英文摘要
Error-penalized scoring rules ($+1$ for a correct answer, $-λ$ for a wrong one, $0$ for abstaining) are increasingly prescribed against hallucination: a rational agent facing such a rule answers exactly when its correctness probability exceeds Chow's threshold $t^\ast=λ/(1+λ)$. We prove that a KL-anchored gradient learner can do the opposite. When abstention is a discrete action, the reward gradient and the anchor's restoring force are throttled by the same gate-saturation factor and die together: under explicit conditions (among them, blanket answering loses score in expectation and prompts share a bounded readout) the model drifts toward refusing everything, its mean training reward rising to zero like $1/t$ in training time $t$, so the curve reads as improvement while coverage collapses. The advantage estimator compounds the failure: in its sparse-answer regime, group normalization silently replaces every designed penalty with an effective penalty of one, moving the learned threshold from $t^\ast$ to $1/2$. The repair is structural: train a mandatory confidence report with a strictly proper score plus a correctness reward, and abstain only at deployment by thresholding the report. The always-emitted report has no gate to saturate, so no shared factor can kill its reward gradient and its anchor together, and its calibrated optimum is attracting. Simulations confirm every prediction, and experiments on language models at two scales confirm the mechanism live: the rule silences questions the models demonstrably still solve within ten optimizer steps, an ablation isolates the cause, and report-level training raises coverage, accuracy, and calibration together.