发表机构
Dartmouth College; University of Pennsylvania; National University of Singapore(达特茅斯学院; 宾夕法尼亚大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM偏好对齐导致的过度自信与校准问题,提出基于双层优化的校准方法,通过最大化预测分布熵实现,在多项选择与开放式问答任务中提升了校准效果与域外泛化能力。
AI 中文摘要
偏好对齐常使大语言模型(LLM)产生过度自信且校准效果差的问题。传统的事后温度缩放具有固有领域依赖性:在某一领域拟合的温度无法跨领域泛化。为此,我们提出在训练期间修改模型参数以提升校准效果。我们将预测分布的熵最大化作为校准目标,通过抑制过于集中的预测直接针对过度自信问题。受温度缩放启发,我们通过双层优化公式实现该目标:下层在参数化损失下训练模型,上层选择损失超参数以最大化熵。为使该框架适用于LLM规模,我们采用高效的一阶近似,避免显式二阶计算。在多项选择及开放式生成问答任务中,实验表明我们的方法可生成校准良好的LLM,且在域外泛化方面具有显著优势。
英文摘要
Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Traditional post-hoc temperature scaling is inherently domain-dependent: a temperature fitted on one domain does not generalize across domains. This motivates us to modify model parameters during training to improve calibration. We propose maximizing the entropy of predictive distributions as the calibration objective, which directly targets overconfidence by discouraging overly concentrated predictions. Inspired by temperature scaling, we realize this through a bilevel optimization formulation, where the lower level trains the model under a parametric loss and the upper level selects loss hyperparameters to maximize entropy. To make the framework practical at LLM scale, we adopt an efficient first-order approximation that avoids explicit second-order computation. Across both multiple-choice and open-ended generative question answering, experiments demonstrate that our method yields well-calibrated LLMs with particular advantages in out-of-domain generalization.
CommentsThird Conference on Language Modeling (COLM 2026)