arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12687cs.CLcs.AI

从批评到置信度:用于基于语言的定量预测和置信度估计的近端策略优化算法

From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation

Mehak Dhaliwal, Rasta Tadayon, Andong Hua, Haewon Jeong, Yao Qin

首次发表
浏览论文内容

中文总结 AI 辅助

研究基于语言的定量预测中模型易出错问题,提出CARE-PPO强化学习框架,通过建立损失预测与近端策略优化微调的联系,联合学习数值估计与置信度信号,在多任务和模型规模上提升预测性能与置信度估计,减少过拟合。

中文摘要 AI 辅助

大语言模型可以对非结构化输入执行基于语言的定量预测,但容易出现幻觉和过度自信的错误,因此了解模型预测内容及其何时可信至关重要。我们引入了CARE-PPO,这是一个强化学习框架,它在用于不确定性估计的损失预测和基于演员-评论家的近端策略优化微调之间建立了联系,能够在基于语言的定量预测中联合学习准确的数值估计和可靠的置信度信号。CARE-PPO使用基于预测误差定义的置信度对齐估计奖励,为演员提供密集的误差感知反馈,同时促使评论家学习与预测质量对齐的价值函数。在推理过程中,我们将评论家重新用作置信度估计器。在医疗保健和金融领域的两项实际任务以及两个Qwen-3模型规模(4B和8B)上,CARE-PPO实现了强大的定量预测性能,同时通过评论家产生了比基于对数几率和语言化基线更好对齐的置信度估计。在跨领域的现实分布外设置下,这些优势依然存在。最后,CARE-PPO减少了在一般指令跟随提示上的任务特定过拟合,这与强化学习微调相对于监督方法的更广泛泛化优势一致。

英文摘要

LLMs can perform language-based quantitative prediction from unstructured inputs, but remain susceptible to hallucinations and overconfident errors, making it critical to know not only what a model predicts, but when its predictions can be trusted. We introduce CARE-PPO, a reinforcement learning framework that establishes a connection between loss prediction for uncertainty estimation and actor-critic PPO fine-tuning, enabling joint learning of accurate numerical estimates and reliable confidence signals in language-based quantitative prediction. CARE-PPO uses a Confidence-Aligned Reward for Estimation, defined as a function of prediction error, to provide dense error-aware feedback to the actor while inducing the critic to learn a value function aligned with prediction quality. During inference, we repurpose the critic as a confidence estimator. Across two real-world tasks in healthcare and finance and two Qwen-3 model scales (4B and 8B), CARE-PPO achieves strong quantitative prediction performance, while producing significantly better-aligned confidence estimates through the critic than logit-based and verbalized baselines. These gains persist under realistic out-of-distribution settings across domains, spanning linguistic and domain shifts. Finally, CARE-PPO reduces task-specific overfitting on general instruction-following prompts, consistent with the broader generalization advantages of RL fine-tuning over supervised approaches.

发表机构

  • University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校)

机构由 AI 辅助整理,请以论文原文为准。

↑