arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

让信用跟随计算:面向大语言模型强化学习的架构感知信用传输

Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning

Qifan Shi, Zhaolu Kang, Chenghua Zhu

arXiv 2608.21501首次发表:更新:

AI 中文总结

该研究针对大语言模型强化学习的信用传输算子与架构无关的问题,提出计算条件信用传输框架及 CompPO 算法,在 Qwen3-4B 等模型上较 GRPO 取得更优的保留准确率与贪心评估表现。

AI 中文摘要

大语言模型强化学习(LLM RL)中的信用分配可分为三类对象:成功证据、将该证据转换为 token 级优势的传输算子,以及将优势转化为策略变化的更新几何。近期研究已在证据、采样和更新几何方面取得显著进展,但传输算子通常与架构无关:固定折扣 GAE 沿 token 时间应用平稳几何核,组相对方法则将结果统计量广播至整个响应,两类算子均未体现 Transformer 策略自身使用的轨迹特定计算。我们提出计算条件信用传输(CCT)这一通用框架,其中行为策略内部计算的分离统计量可参数化因果核,用于将下游价值通过 rollout 传输。具体算法 CompPO 将原生注意力集中度映射为有界的逐 token 保留门,该门同时用于单步自举和依赖路径的广义优势迹(Comp-GAE),并协同设计了传输对齐评论家(TAC),其复用 Actor 的隐藏状态和路由信息,无需第二个同等规模的 Transformer。任务奖励和裁剪后的 PPO 策略目标保持不变,常数门可恢复固定系数 GAE。在 5 个 Qwen3-4B 随机种子上,CompPO 达到 61.4% 的最终保留准确率(95% 置信区间 [60.8,62.0]),而调优后的 GRPO 为 53.8%([52.9,54.7]);采用标准评论家的 Comp-GAE(55.2%)和采用固定门的 TAC(56.4%)均不及完整模型,交互增益为 2.4([1.9,2.9])。打乱和位置控制实验证实了轨迹特定对齐,CompPO 在 12 次 PPO 网格运行中 10 次稳定,而 GRPO 仅 3 次;在 Qwen3-4B 和 Llama-3.1-8B-Instruct 上,冻结评估较 GRPO 分别提升 4.3 和 3.9 个贪心 pass@1 宏观点。

英文摘要

Credit assignment in large-language-model reinforcement learning (LLM RL) can be separated into three objects: evidence about success, a transport operator that converts this evidence into token-level advantages, and an update geometry that turns advantages into policy changes. Recent work has greatly improved evidence, sampling, and update geometry, but the transport operator is usually architecture-agnostic. Fixed-discount GAE applies a stationary geometric kernel along token time; group-relative methods broadcast an outcome statistic across an entire response. Neither operator represents the trajectory-specific computation used by the Transformer policy itself. We introduce computation-conditioned credit transport (CCT), a general framework in which a detached statistic of the behavior policy's internal computation parameterizes the causal kernel that transports downstream value through a rollout. Our concrete algorithm, CompPO, maps native attention concentration to a bounded per-token retention gate, uses the gate in both the one-step bootstrap and a path-dependent generalized-advantage trace (Comp-GAE), and co-designs a transport-aligned critic (TAC) that reuses the actor's hidden states and routing information without a second same-scale Transformer. The task reward and clipped PPO policy objective remain unchanged; a constant gate recovers fixed-coefficient GAE. Across five Qwen3-4B seeds, CompPO reaches 61.4% final held-out accuracy (95% CI [60.8,62.0]) versus 53.8% [52.9,54.7] for tuned GRPO. Neither Comp-GAE with a standard critic (55.2%) nor TAC with a fixed gate (56.4%) matches the full model (interaction +2.4 [1.9,2.9]). Shuffle and position controls confirm trajectory-specific alignment; CompPO is stable in 10/12 PPO-grid runs versus 3/12. Frozen evaluation improves over GRPO by 4.3 and 3.9 greedy pass@1 macro points on Qwen3-4B and Llama-3.1-8B-Instruct.

Comments22 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑