arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

融合Gromov-Wasserstein几何上的RLVR分层信用分配

Hierarchical Credit Assignment for RLVR on Fused Gromov-Wasserstein Geometry

Qi Yu, Ruizhong Qiu, Zhichen Zeng, Xuying Ning, Yanjun Zhao, Dongqi Fu, Yinglong Xia, Hong Li, Hanghang Tong

arXiv 2610.04344首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Meta(伊利诺伊大学厄巴纳-香槟分校; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出HarA,一种基于FGW几何的分层信用分配方法,通过语义新颖性信号重加权token优势,增强组级RLVR中LLM的探索,并在多种推理基准上超越现有方法。

AI 中文摘要

基于可验证奖励的强化学习(RLVR)已被证明能够提升大语言模型(LLMs)在多种推理任务中的推理能力。然而,基于组的RLVR方法(如GRPO)对同一结果内所有rollout中的token赋予统一的优势值。虽然现有工作基于局部信号(如token位置或熵)对GRPO的信用分配进行细化,但这些方法往往无法捕捉推理行为相对于当前策略的全局语义新颖性。在本工作中,我们提出了一种针对基于组的RLVR方法的分层信用分配方法,称为HarA,该方法在RLVR过程中识别并鼓励语义新颖的推理行为。HarA将每个采样的rollout表示为隐藏状态和token位置的分布,并计算具有相同结果的所有rollout的融合Gromov-Wasserstein(FGW)重心,从而在当前策略下捕捉潜在空间中的内部推理模式。推理元素的语义新颖性可以通过其对当前rollout与重心之间FGW距离的贡献来衡量。虽然求解FGW公式计算成本高昂,但我们引入了一种锚引导线性化方法,将其转化为可通过Sinkhorn算法高效求解的Wasserstein公式。通过基于新颖性信号重新加权基于组的RLVR方法的token级优势,HarA以灵活的粒度突出新颖推理行为,以促进LLM的细粒度探索。在三种基于组的RLVR方法上的大量实验表明,我们的即插即用方法有效增强了LLM的探索能力,在多种推理基准上优于现有方法。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens within rollouts of the same outcome. While existing works refine credit assignment of GRPO based on local signals such as token locations or entropy, they often fail to capture the global semantic novelty of a reasoning behavior relative to the current policy. In this work, we propose a hierarchical credit assignment approach for group-based RLVR methods, called HarA, which identifies and encourages semantically novel reasoning behaviors during RLVR. HarA represents each sampled rollout as a distribution over the hidden states and locations of tokens, and computes the Fused Gromov-Wasserstein (FGW) barycenters of all rollouts with the same outcome, capturing the internal reasoning patterns in the latent space under the current policy. The semantic novelty of a reasoning element can then be measured by its contribution to the FGW distance between the current rollout and the barycenter. While solving the FGW formulation is expensive, we introduce an anchor-guided linearization that turns it into a Wasserstein formulation solvable via the Sinkhorn algorithm efficiently. By reweighing token-level advantage of group-based RLVR methods based on the novelty signals, HarA highlights novel reasoning behaviors at flexible granularities to encourage fine-grained LLM exploration. Extensive experiments across three group-based RLVR methods show that our plug-and-play method effectively enhances the exploration of LLMs, outperforming existing methods across diverse reasoning benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑