发表机构
School of Computing National University of Singapore; Institute of Advanced Intelligence and Computing; Agency for Science, Technology and Research; Department of Electrical & Computer Engineering(新加坡国立大学计算机学院; 高级智能与计算研究所; 科技研究局; 电气与计算机工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型用强化学习训练以提升推理准确性和表达置信度时,奖励函数的设计问题。提出不可破解置信奖励方案概念并定义相关奖励方案谱,指出实际数据集中存在的问题及最佳校准奖励方案因数据集和应用而异。
AI 中文摘要
本文考虑大语言模型使用强化学习训练以同时提高推理准确性并表达其置信度的情况。我们的奖励方案使用两个函数奖励大语言模型表达的置信度:大语言模型正确时一个,错误时另一个。设计不佳的奖励方案可能导致大语言模型被激励给出错误答案以使其能自信其答案确实错误,即置信奖励破解现象。我们提出不可破解置信奖励方案的概念并为大语言模型中的强化学习置信度校准训练定义此类奖励方案谱。我们证明在实际数据集中,使用未设计为不可破解的奖励方案时会发生选择性置信奖励破解。我们还证明与准确性权衡校准最佳的奖励方案取决于数据集和应用,并建议将奖励方案用作超参数以根据应用的重要性优化权衡。我们实验的代码可在这个https网址获取。
英文摘要
In this paper, we consider the setting where large language models (LLMs) are trained using reinforcement learning (RL) to simultaneously improve reasoning accuracy and verbalize their confidence. Our reward scheme uses two functions for rewarding confidence verbalized by the LLM: one for correct answers and the other for incorrect answers. If poorly designed, such a scheme may incentivize an LLM to answer incorrectly in order for its confidence to be calibrated, a phenomenon we term confidence reward hacking. We introduce the notion of non-hackable confidence reward schemes and provide methods for constructing them. We show that selective confidence reward hacking can arise in practical datasets under hackable reward schemes while non-hackable reward schemes are resistant to hacking. Finally, we place some of these schemes along an overconfidence-underconfidence spectrum for RL-based confidence calibration and demonstrate experimentally that they tend to exhibit the corresponding calibration biases relative to other schemes in the spectrum. The code of our experiments is available in https://anonymous.4open.science/r/rl-confidence-calibration-9ED4/README.md.
Comments80 pages, 10 figures