arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28771cs.LG

ERR+: 用于高效且果断的LLM推理的顺序熵分辨率

ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning

Xin Jiang, Minhao Wang, Wen Wu, Zhentao Xie, Shangheng Du, Jinxin Shi, Jiabao Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对RLVR方法对推理过程质量指导不足的问题,提出两阶段RLVR框架ERR+,通过顺序优化熵缓解奖励与稳健相对效率奖励,在五个数据集上实现了LLM推理准确率与简洁性的提升。

中文摘要 AI 辅助

大型推理模型通过使用可验证奖励的强化学习(RLVR)生成扩展的思维链(CoT)轨迹,在复杂任务上实现了强大性能。尽管当前RLVR方法基于正确性的奖励信号取得了优异结果,但它们对推理过程本身的质量指导有限,使得内部推理结构基本未优化。通过对多个模型系列的实证分析,我们发现了一个一致模式:正确的推理轨迹在思维阶段的 token 级熵下降比错误轨迹更频繁、幅度更大。我们提出ERR+,一个基于该观察的两阶段RLVR框架。第一阶段使用熵缓解奖励(ERR)训练,该奖励与思维阶段累积的token级熵下降成比例,并按响应长度进行对数归一化。与之前抑制熵的方法不同,ERR奖励不确定性的解决,同时不约束探索性的高熵状态。第二阶段引入稳健相对效率奖励,通过组内z-score的tanh变换对每个响应的长度相对于共同生成的同类响应进行评分。我们提供的形式分析表明,两个目标的联合优化在训练早期会引发梯度冲突,从而推动了顺序设计。在五个数据集上的实验表明,在不同模型主干上,准确率和响应简洁性均实现了持续提升。我们的代码可在此处获取:this https URL

英文摘要

Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR). While current RLVR methods have achieved strong results with correctness-based reward signals, they provide limited guidance on the quality of the reasoning process itself, leaving the internal reasoning structure largely unoptimized. Through empirical analysis across multiple model families, we identify a consistent pattern: correct reasoning trac es exhibit more frequent and larger token-level entropy drops within the thinking phase than incorrect ones. We propose ERR+, a two-phase RLVR framework grounded in this observation. The first phase trains with the Entropy Relief Reward (ERR), a bonus proportional to cumulative token-level entropy drops in the thinking phase, log-normalized by response length. Unlike prior methods that suppress entropy, ERR rewards the resolution of uncertainty while leaving exploratory high-entropy states unconstrained. The second phase introduces the Robust Relative Efficiency Reward, which scores each response's length against co-generated peers via a $\tanh$-transformed within-group $z$-score. We provide a formal analysis showing that joint optimization of the two objectives induces gradient conflict in early training, motivating the sequential design . Experiments on five datasets demonstrate consistent improvements in both accuracy and response conciseness across model backbones. Our code is available at https://github.com/XrkArul/err_response

发表机构

  • School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)
  • ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑