发表机构
Eastern Institute of Technology; The Hong Kong Polytechnic University(东理工学院; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大型语言模型训练后方法在推理中重塑置信度的情况,提出三阶段校准框架,发现不同方法在不同阶段的置信度特点,基于此提出位置感知置信策略PosConf,提升了强化学习答案聚合及在线策略蒸馏早期停止效果。
AI 中文摘要
大型语言模型通过监督微调、强化学习和在线策略蒸馏在推理方面取得了显著进展,但这些训练后方法通常仅通过最终答案准确性进行评估。我们研究它们在推理过程中如何重塑置信度。我们引入了一个三阶段校准框架,在思维链生成之前、期间和之后评估置信度,分别对应难度估计、早期终止和答案聚合。通过在数学推理基准上的对比,发现在线策略蒸馏提供最有用的推理前置信度,监督微调给出最强的早期停止在线信号,强化学习产生最可靠的聚合跟踪级信号。我们进一步表明置信度可靠性与位置有关:强化学习置信度在路径承诺阶段后变得信息丰富,而在线策略蒸馏置信度早期有用但后期可能反向校准。基于此,我们提出了PosConf,一种位置感知置信策略,仅使用可靠相对位置区间的置信度。PosConf在多数投票基础上将强化学习答案聚合提高了6.1分,在严格令牌预算下持续改进在线策略蒸馏早期停止,通过避免其后期反向校准区域获得高达4.3分的提升,表明推理模型中的置信度应按阶段和位置感知方式使用。
英文摘要
Large language models have made strong reasoning gains through supervised fine-tuning, reinforcement learning, and on-policy distillation, yet these post-training methods are usually evaluated only by final-answer accuracy. We study how they reshape confidence during reasoning. We introduce a three-stage calibration framework that evaluates confidence before, during, and after chain-of-thought generation, corresponding to difficulty estimation, early termination, and answer aggregation. Through a controlled comparison on mathematical reasoning benchmarks, we find that OPD provides the most useful pre-reasoning confidence, SFT gives the strongest online signal for early stopping, and RL produces the most reliable trace-level signal for aggregation. We further show that confidence reliability is position-dependent: RL confidence becomes informative after a path-commitment phase, while OPD confidence is useful early but can become inversely calibrated later. Based on this observation, we propose PosConf, a position-aware confidence strategy that uses confidence only from reliable relative-position intervals. PosConf improves RL answer aggregation by 6.1 points over majority voting and consistently improves OPD early stopping under tight token budgets, with gains up to 4.3 points by avoiding its later inverse-calibration region, showing that \emph{confidence in reasoning models should be used both stage-wise and position-awarely}. Our code is available at https://github.com/EIT-NLP/Post-Training-Calibration.