arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

恰当评分规则如何塑造大语言模型(LLM)的预测

How Proper Scoring Rules Shape LLM Forecasting

Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satopää, Philip E. Tetlock

arXiv 2608.28482首次发表:更新:

发表机构

Lightning Rod Labs; INSEAD; School of Arts & Sciences, University of Pennsylvania(闪电杆实验室; 欧洲工商管理学院; 宾夕法尼亚大学文理学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究了5种恰当评分规则作为训练目标对LLM预测的影响,发现不同规则训练的模型在校准、概率使用等方面有差异,Brier训练模型表现最优。

AI 中文摘要

本文评估了奖励函数的选择如何影响大语言模型(LLM)预测器的性能与行为。我们将5种恰当评分规则作为训练目标,用于对已解决的现实世界事件进行二元预测。尽管这些规则在理论上对如实概率报告具有相同的激励作用,但由此得到的模型在校准程度、概率使用方式以及偏差、信息和噪声的估计分布上存在差异,而在总体准确率和区分度上的差异较小。经Brier训练的模型具有最低的观测Brier分数和最高的AUC-ROC,而经log训练的模型具有最高的观测log分数和最低的校准误差。具有相似总体性能的模型也通过偏差、信息和噪声的不同组合达到该性能。因此,恰当评分规则作为训练目标并非可互换的,奖励函数的选择不仅可能影响LLM预测的效果,还可能影响其预测误差的构成。每个条件使用单个随机种子,因此部分差异可能反映训练随机性。

英文摘要

This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same theoretical incentive for truthful probability reporting, the resulting models differ in calibration, probability use, and estimated profiles of bias, information, and noise, with smaller differences in aggregate accuracy and discrimination. The Brier-trained model has the lowest observed Brier score and highest AUC-ROC, while the log-trained model has the highest observed log score and lowest calibration error. Models with similar aggregate performance also reach that performance through different combinations of bias, information, and noise. Proper scoring rules therefore need not behave interchangeably as training objectives. Reward choice may shape not only how well an LLM forecasts, but how its forecasting errors are structured.

CommentsAdded a three-seed replication of the Log and Brier reward conditions, including forecasting metrics and BIN analysis

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑