arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

揪出搭便车者:基于夏普利值的强化学习并行推理奖励归因

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning

Wentao Zhang, Haoyu Zhang, Xinke Jiang, Yuxuan Cheng, Yuhan Pan, Miao Li, Zhipeng Qiao, Tao Feng, Zhen Tao, Dengji Zhao

arXiv 2607.18979首次发表:更新:

发表机构

ShanghaiTech University; City University of Hong Kong; Peking University; The Chinese University of Hong Kong, Shenzhen; Zhejiang University(上海科技大学; 香港城市大学; 北京大学; 香港中文大学(深圳); 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大型语言模型多步推理中并行推理路径贡献难区分问题,提出基于夏普利值的强化学习框架并行夏普利值方法,通过量化边际贡献等实现按比例分配奖励,实验证明其优于基线且能改进多路径推理。

AI 中文摘要

大型语言模型在多步推理方面表现出色,但当前并行推理方法往往无法区分各推理路径的贡献。许多路径可能冗余、误导甚至有害,而结果级奖励分配均匀,导致学习信号模糊和训练不稳定。我们提出了并行夏普利值方法,这是一种强化学习框架,可在多路径推理中对细粒度的路径级贡献进行归因。将每条路径视为合作博弈中的参与者,利用夏普利值量化边际贡献,使用生成奖励模型评估路径效用,并通过蒙特卡罗采样进行有效近似。在数学推理基准上的实验表明,并行夏普利值方法优于现有基线,同时提供更稳定和可解释的训练。我们的框架有效地“揪出搭便车者”,按比例分配奖励并改进大型语言模型中的多路径推理。

英文摘要

Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contributions of individual reasoning paths. Many paths may be redundant, misleading, or even detrimental, but outcome-level rewards assign uniform reward, leading to ambiguous learning signals and unstable training. We propose Parallel Shapley, a reinforcement learning framework that attributes fine-grained, path-level contributions in multi-path reasoning. Treating each path as a player in a cooperative game, we leverage Shapley values to quantify marginal contributions, using a generative reward model to evaluate path utilities and Monte Carlo sampling for efficient approximation. Experiments on mathematical reasoning benchmarks show that Parallel Shapley outperforms existing baselines while providing more stable and interpretable training. Our framework effectively "fishes out the free riders," assigning reward proportionally and improving multi-path reasoning in LLMs.

Comments19 pages, 4 figures, 8 Tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑