arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38850cs.AIcs.CL

OpenJev-RLCD:一个可运行的RLCD实现

OpenJev-RLCD: A Working RLCD Implementation

Zhimin Gao, Pichao Wang

AI总结:

本文提出RLCD(校准决策强化学习)的可运行实现,通过两阶段(先校准后强化)解决推理模型过度自信问题,在准确率和选择性预测上优于SFT、RFT/STaR和GRPO,并证明在标注分歧场景下无法超越交叉熵。

AI中文摘要:

诸如Jev之类的决策模型以概率形式回答问题,而这些概率只有在经过校准后才具有实用性。开源复现依赖于监督微调加温度缩放,而基于可验证奖励的强化学习(RLVR)则使推理模型过度自信。我们提出了一种针对推理模型的校准决策强化学习(RLCD)的可运行实现:模型采样一个推理链,随后我们使用严格适当的评分规则对其承诺的答案分布进行评分。一个方差恒等式表明,对多个样本的混合进行评分会奖励不一致的推理链,而RLVR正是这个混合目标减去其多样性项。若进行朴素优化,每个推理链的目标要么关闭推理,要么被策略梯度噪声淹没,这导致了一个两阶段方案:先校准,再强化。使用Qwen3-1.7B在两个推理任务上(3个种子,配对测试),RLCD在准确率上匹配或优于SFT、RFT/STaR和GRPO(各自经过温度缩放),并在选择性预测上优于所有方法;在GSM8K答案验证中,单次查询在误差≤5%的情况下决定了\\(\gvTwoCovFive\\%\\)的项目,而GRPO为\\(\gvGrpoCovFive\\%\\)。当不确定性来源于标注者分歧时,RLCD可证明无法超越交叉熵。代码和结果:此https URL。

英文摘要:

Decision models such as Jev answer questions with probabilities, which are only useful if they are calibrated. Open-source reproductions rely on supervised fine-tuning plus temperature scaling, while reinforcement learning from verifiable rewards (RLVR) makes reasoning models overconfident. We present a working implementation of reinforcement learning for calibrated decisions (RLCD) for reasoning models: the model samples a rationale, and we score the answer distribution it commits to afterwards with a strictly proper scoring rule. A variance identity shows that scoring the mixture of several samples rewards disagreeing rationales, and that RLVR is exactly this mixture objective without its diversity term. Optimized naively, the per-rationale objective either switches reasoning off or is drowned out by policy-gradient noise, which leads to a two-stage recipe: calibrate, then reinforce. With Qwen3-1.7B on two reasoning tasks (3 seeds, paired tests), RLCD matches or beats SFT, RFT/STaR and GRPO (each temperature-scaled) in accuracy and beats all of them in selective prediction; on GSM8K answer verification a single query decides \gvTwoCovFive\% of the items at $\le$5\% error, versus \gvGrpoCovFive\% for GRPO. When uncertainty comes from annotator disagreement, RLCD provably cannot beat cross-entropy. Code and results: https://github.com/ZimmyGao/openjev-rlcd.

↑