arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32444cs.LGcs.AI

重新思考大语言模型强化学习中的训练-推理不匹配:其来源与纠正方法

Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

Tianrun Yu, Kaixiang Zhao, Shangzhe Li, Yuxiao Yang, Porter Jenkins, Weitong Zhang, Taylor W. Killian

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型强化学习中训练与推理引擎概率不一致的问题,提出校准重要性采样(CIS),通过置信度感知截断修正策略更新,在三个混合专家模型和五个数学基准上取得最佳平均性能。

中文摘要 AI 辅助

我们研究了大语言模型在可验证奖励强化学习(RLVR)中的训练-推理不匹配问题,其中轨迹由推理引擎采样,而梯度由训练引擎计算,两个引擎对相同令牌赋予不同的概率。为了在策略更新中考虑这种差异,我们引入了校准重要性采样(CIS)。CIS 的动机是一个经验支持的 logit 位移表征,该表征将不匹配表示为 log-odds 中的加性位移 $\varepsilon_t$,由 softmax 之前的每个 logit 扰动决定,其分布近似于对令牌置信度不变。这一表征激发了置信度感知的截断:大的正位移在单一常数阈值处截断,这映射为随令牌置信度增加而收紧的重要性比率上限。理论上,我们证明 CIS 将控制精确重要性采样误差的无界二阶矩替换为常数有界的项,代价是由截断超额控制的偏差。在三个混合专家模型和五个数学推理基准的评估中,CIS 在所有三个模型上取得了五个基准的平均最高分,优于评估的基线。诊断分析表明,与截断重要性采样相比,CIS 对低置信度令牌施加的截断偏差更小,而向上裁剪小重要性权重会降低留出准确率。

英文摘要

We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To account for this discrepancy in policy updates, we introduce calibrated importance sampling (CIS). CIS is motivated by an empirically supported logit-displacement characterization that expresses the mismatch as an additive displacement $\varepsilon_t$ in log-odds, determined by the per-logit perturbation before the softmax, whose distribution is approximately invariant to token confidence. This characterization motivates a confidence-aware truncation: large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases. Theoretically, we show that CIS replaces the unbounded second moment that governs the error of exact importance sampling with a term bounded by a constant, at the cost of a bias controlled by the truncated excess. In evaluation across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines. Diagnostic analyses show that CIS places less truncation bias on low-confidence tokens than truncated importance sampling, while upward clipping of small importance weights reduces held-out accuracy.

补充信息

↑