arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推进RLVR中的熵级信用分配:基于邻近熵策略优化

Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization

Yun Kim, Nojun Kwak

arXiv 2609.39402首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出邻近熵策略优化(PEPO),通过局部上下文度量token重要性,改进RLVR中的信用分配,在数学推理任务上超越GRPO和熵基线。

AI 中文摘要

无值模型的RLVR方法(如GRPO)在rollout中为所有token分配均匀优势,忽略了token贡献的不均等性。近期方法使用token熵作为重要性代理,但全局计算熵,混淆了重要性与提示难度和位置趋势。我们认为重要性应相对于每个token的局部上下文来衡量。我们引入邻近熵,一种相对于邻近token的局部token重要性度量,并证明其对两种混淆因素不变。邻近熵策略优化(PEPO)利用该度量加权每个token的优势,在Qwen3-1.7B、Qwen3-4B和Llama-3.2-3B-Instruct上的数学推理任务中优于GRPO和基于熵的基线。我们还展示了该公式可推广到其他算法,将邻近熵替换到现有方法中可提升性能,且应用于单流RL时,在全局熵失败的情况下取得成功。

英文摘要

Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to the local context of each token. We introduce proximal entropy, a local measure of token importance relative to neighboring tokens, and prove it is invariant to both confounders. Proximal Entropy Policy Optimization (PEPO) uses it to weight per-token advantages and outperforms GRPO and entropy-based baselines on mathematical reasoning across Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct. We also show the formulation generalizes to other algorithms where substituting proximal entropy into existing methods improves, and applying it to single-stream RL succeeds where global entropy fails.

Comments21 pages, 4 figures. Accepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑