arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

给予应有的功劳:冗余感知学习用于高效推理

Giving Credit Where It's Due: Redundancy-Aware Learning for Efficient Reasoning

Yuqing Zhou, Hong Wang, Manqing Mao, Zhuoer Wang, Samson Koelle, Jie Yuan, Yanjun Lin, James Feng, Nikki Lijing Kuang, Ziwei Zhu, Wei Niu

arXiv 2609.27156首次发表:更新:

发表机构

George Mason University; Amazon, Inc(乔治梅森大学; 亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大型推理模型冗长推理问题,提出冗余感知信用分配方法RECAP,通过结构责任和步骤效能分配信用,无需额外奖励模型,在两个7B模型和四个数学基准上提升了准确率-效率权衡。

AI 中文摘要

大型推理模型能够产生正确但冗长的推理轨迹。现有方法通过轨迹级目标或局部token和步骤级信号来提高推理效率,但很少对步骤间的语义依赖进行建模。这限制了它们区分冗余步骤与支持后续推理步骤的能力,使得在不牺牲准确性的情况下缩短推理变得更加困难。我们引入了RECAP(通过传播进行冗余感知的信用分配),该方法通过基于步骤在推理结构中的下游角色及其对正确解决问题的贡献来分配应有的功劳,从而解决了这一局限。我们定义了结构责任来捕捉步骤的下游角色,通过衡量后续推理对其依赖的强度,利用从最终答案节点通过一个与结果无关、由LLM标注的语义依赖图反向传播的信用来实现。然而,一个步骤可能具有高结构责任,却将推理引向远离正确答案的方向。因此,RECAP引入了步骤效能,通过每一步添加时金标准答案对数似然的变化来衡量对答案导向的进展。这些信号共同将rollout级别的GRPO优势重塑为步骤特定的更新。RECAP既不需要单独训练的过程奖励模型,也不需要预先构建的简洁轨迹。在两个7B模型和四个数学推理基准上,RECAP改善了准确率-效率的权衡。在Qwen2.5-Math-7B上,相对于GRPO,它在所有四个基准上将pass@1提高了2.0-3.7个百分点,同时将推理token减少了8%-31%。分析表明,这些节省反映了更少的推理操作和更少的死胡同推理,而非更紧凑的表达。

英文摘要

Large reasoning models can produce correct yet unnecessarily long reasoning traces. Existing methods improve reasoning efficiency with trajectory-level objectives or local token- and step-level signals, but rarely model inter-step semantic dependencies. This limits their ability to distinguish redundant steps from those that support later deductions, making it harder to shorten reasoning without sacrificing accuracy. We introduce RECAP (REdundancy-aware Credit Assignment via Propagation), which addresses this limitation by assigning credit where it is due based on both a step's downstream role in the reasoning structure and its contribution to solving the problem correctly. We define structural responsibility to capture the step's downstream role by measuring how strongly later reasoning depends on it, using credit propagated backward from the final-answer node through an outcome-independent, LLM-annotated semantic dependency graph. However, a step can have high structural responsibility yet steer the reasoning away from the correct solution. RECAP therefore introduces step efficacy to measure answer-directed progress through changes in gold-answer log-likelihood as each step is added. Together, these signals reshape rollout-level GRPO advantages into step-specific updates. RECAP requires neither a separately trained process reward model nor preconstructed concise trajectories. Across two 7B models and four mathematical reasoning benchmarks, RECAP improves the accuracy-efficiency trade-off. On Qwen2.5-Math-7B, it improves pass@1 by 2.0-3.7 percentage points while reducing reasoning tokens by 8%-31% relative to GRPO across all four benchmarks. Analysis suggests these savings reflect fewer reasoning operations and less dead-end reasoning, rather than more compact expression.

Comments26 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑