发表机构
Shanghai Key Laboratory of Scalable Computing and Systems; School of Computer Science, Shanghai Jiao Tong University(上海市可扩展计算与系统重点实验室; 上海交通大学计算机科学与工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM推理长度错配问题,提出CARE方法,通过对比准确率奖励估计实现自适应长度奖励,在多个基准上提升Pass@1达4%并减少37%推理长度。
AI 中文摘要
强化学习(RL)已被证明能有效提升大型语言模型(LLMs)的推理性能,尤其是在复杂的数学和编程任务中。然而,这种能力伴随着系统性的长度错配问题,即模型在简单问题上过度推理,而在较难问题上过早终止,从而降低推理效率且准确率提升甚微。许多长度自适应方法通过根据问题难度分配令牌预算来缓解这一问题,其隐含假设是较难的问题从更长的推理中单调受益。相比之下,我们发现推理长度对准确率的影响集中在部分可解的问题上。进一步分析揭示,显式的长度奖励可能产生非预期的训练动态。基于这些发现,我们提出CARE——对比准确率奖励估计——它比较在线采样响应中每个问题的有益长度调整,并在组相对策略优化中应用自适应长度奖励,无需额外超参数或推理成本。在多个推理基准上的实验表明,我们的方法将Pass@1提高了最多4%,同时将推理长度减少了37%,实现了更高的令牌效率。代码将在论文被接受后公开。
英文摘要
Reinforcement learning (RL) has proven effective in enhancing the reasoning performance of large language models (LLMs), particularly in complex mathematical and programming tasks. However, this capability comes with systematic \textit{length misallocation}, in which models devote excessive reasoning to simple questions while terminating prematurely on harder ones, degrading inference efficiency with negligible accuracy improvement. Many length-adaptive methods mitigate this issue by allocating token budgets according to question difficulty, under the implicit assumption that harder questions benefit monotonically from extended reasoning. In contrast, we find that the effect of reasoning length on accuracy is concentrated on \textit{partially solvable} questions. Our further analysis reveals that explicit length rewards can produce unintended training dynamics. Motivated by these findings, we propose \textbf{CARE}---\textbf{C}ontrastive \textbf{A}ccuracy \textbf{R}eward \textbf{E}stimation---which compares the beneficial length adjustment per question from online sampled responses and applies adaptive length rewards within Group Relative Policy Optimization, with no extra hyperparameters or additional inference cost. Experiments across multiple reasoning benchmarks demonstrate that our method improves Pass@1 by up to \(4\%\) while simultaneously reducing reasoning length by \(37\%\), achieving higher token efficiency. Code will be available upon the acceptance of this paper.