arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30572cs.LGstat.ML

熵正则化:针对已验证示范的交叉熵的免费修正

Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

Mihir Dhanakshirur, Adam Ousherovitch, Ambuj Tewari

AI总结:

本文发现交叉熵最小化在验证任务中可能与验证器风险不一致,提出熵正则化交叉熵(ER-CE)作为修正,通过在数学推理和代码生成基准上实验证明其能提高验证器准确性。

AI中文摘要:

大型语言模型通常在专家示范上使用交叉熵(CE)进行后训练,即使下游目标不是模仿示范的解决方案,而是产生任何被验证器接受的输出。这种不匹配在具有多个正确解的验证领域中出现,例如数学推理和代码生成,其中训练数据可能每个问题只包含一个专家解。我们表明,最小化交叉熵可能与最小化验证器风险不一致;两个策略可以对观察到的示范赋予相同的似然,同时将不同的概率质量分配给不正确的输出。这通过一个学习理论的反例形式化,其中CE最小化选择了一个次优策略。我们确定,控制学习策略的支持集可以通过防止概率质量扩散到不支持输出来解决这个问题。由于支持集大小不可微且计算上难以处理,我们提出了熵正则化交叉熵(ER-CE),使用词元级香农熵作为可处理的代理。最后,在数学推理和代码生成基准上,我们发现熵正则化训练一致地提高了验证器准确性,优于标准交叉熵。我们的结果识别了在验证任务中基于模仿的后训练的一个简单失败模式,并提供了一种与产生正确输出更一致的实际目标。

英文摘要:

Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematical reasoning and code generation, where training data may contain only one expert solution per problem. We show that minimizing cross-entropy can be misaligned with minimizing verifier risk; two policies can assign identical likelihood to the observed demonstrations while placing different probability mass on incorrect outputs. This is formalized through a learning-theoretic counterexample in which CE minimization selects a suboptimal policy. We identify that controlling the support of the learned policy can solve this problem by preventing probability mass from spreading to unsupported outputs. Since support size is non-differentiable and computationally intractable, we propose entropy-regularized cross-entropy (ER-CE), using token-level Shannon entropy as a tractable proxy. Finally, across mathematical reasoning and code-generation benchmarks, we find that entropy-regularized training consistently improves verifier accuracy over standard cross-entropy. Our results identify a simple failure mode of imitation-based post-training in verifiable tasks and provide a practical objective that is better aligned with producing correct outputs.

↑