arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DiDPO:用于编码智能体训练的差异内差异策略优化

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training

Xucong Wang, Zhe Zhao, Liheng Yu, Di Wu, Xiaofeng Cao, Pengkun Wang

arXiv 2608.07147首次发表:更新:

发表机构

University of Science and Technology of China (USTC); Stanford University; Suzhou Institute for Advanced Research, USTC; Tongji University(中国科学技术大学; 斯坦福大学; 中国科学技术大学苏州高等研究院; 同济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对编码智能体训练的细粒度信用分配难题,提出DiDPO方法,在Qwen2.5-7B-Coder上性能超可比方法10%以上,开源了支持多种RL方法的verl-code代码库。

AI 中文摘要

带可验证奖励的强化学习(RLVR)已成为训练编码智能体的强大范式,其中编译和测试的执行反馈提供了客观验证。然而,与普通智能体任务不同,编码智能体面临独特且更细粒度的信用分配挑战:每一步的编码动作会同时在代码版本的不同区域进行不同修改,导致独立修改的贡献难以区分。现有RLVR方法大多利用结果奖励或步骤级奖励,无法深入代码差异,使编码动作的独特特性在训练中不可见。本文提出无评论者的RL方法DiDPO,直接从代码差异结构构建细粒度信用单元。DiDPO将多轮编码交互组织为多个思考-动作步骤,在采样轨迹中发现代码差异;随后通过“可分组性分数”将每个完整差异拆分为高度相似的子差异,以此选择锚点,该分数提供的拆分方案可最优平衡锚点的语义范围及其可能形成的组规模;最后这些锚点形成优势组,将差异级优势映射回单个响应标记。在长视野编码与推理基准上的实验显示,DiDPO显著优于强大的智能体RL基线;在Qwen2.5-7B-Coder上,DiDPO比可比方法超出10%以上,缩小了与更大规模模型的差距,为编码智能体训练中的细粒度信用分配提供了原则性框架;同时我们开源了verl-code,这是支持多种RL方法和编码基准的智能体RL代码库。

英文摘要

Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides objective verification. However, unlike agent tasks, coding agents face a unique and finer-grained credit assignment challenge: at each step, coding actions simultaneously pack varying changes into different regions of a code version, which makes the contribution of independent change indistinguishable. Existing RLVR methods mostly leverage the outcome reward or step-level reward, which fails to dive into a code diff and makes unique properties of coding actions invisible to training. In this paper, we propose Diff-in-Diff Policy Optimization (DiDPO), a critic-free RL method that constructs fine-grained credit units directly from the structure of code diffs. DiDPO organizes multi-turn coding interactions into multiple thought--action steps and discovers code diffs across sampled trajectories. It then selects anchors by aggregating highly similar sub-diffs split from each whole diff by our ``groupability score'', which provides the splitting schema that optimally balances the semantic scope of anchors and the group mass they may form. Finally these anchors form advantage groups and project the diff-level advantage back to individual response tokens. Experiments on long-horizon coding and reasoning benchmarks show that DiDPO significantly outperforms strong agentic RL baselines. On Qwen2.5-7B-Coder, DiDPO exceeds comparable methods by over 10\% and narrows the gap with far larger models, offering a principled framework for fine-grained credit assignment in coding agent training. We also open-source verl-code, an agentic rl codebase that supports various RL methods and coding benchmarks.

Comments16 pages, 6 figures, work in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑