arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17393cs.AI

LEGO-RL:利用原生强化学习实现编码智能体

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

  • Huawei Technologies Co., Ltd(华为技术有限公司)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai

AI总结:

LEGO-RL 是桥接原生编码智能体 harness 与策略梯度优化的框架,通过三大支柱解决训练不匹配问题,提升 Qwen3.5-35B-A3B 在三个编码基准上的性能,且保持高概率相关性。

AI中文摘要:

编码智能体的强化学习越来越依赖长期运行的智能体 harness 来管理工具集成、仓库上下文和执行反馈。然而,这些 harness 的原生执行环境与策略梯度训练存在固有不匹配:环境崩溃和奖励黑客行为会破坏结果信号,而训练-推理差异会将 rollout 行为与策略更新解耦。为解决这一问题,我们提出 LEGO-RL,这一框架可在不修改内部控制流的情况下,将原生编码智能体 harness 与可扩展的策略梯度优化桥接。LEGO-RL 基于三大支柱构建:(1)通过进程内 LLM 代理实现忠实优化,该代理可捕获原始生成流以实现 token 级对齐,即便在 harness 侧压缩或重新序列化的情况下,也能在训练器侧重新计算稳健的对数概率;(2)通过具备镜像缓存和阶段式防御的可扩展沙箱编排实现可靠执行,以缓解奖励黑客行为;(3)通过集成插件实现可观测训练,该插件可自动执行验证和监控,并搭配用于精细轨迹诊断的 Live UI。我们通过在三个原生编码智能体 harness 上使用 GSPO 训练稀疏 MoE 模型 Qwen3.5-35B-A3B 来评估 LEGO-RL。LEGO-RL 在 SWE-bench Verified 上分别将 Qwen3.5-35B-A3B 在 OpenHands SDK(64.0% 提升至 70.4%)、Claude Code(62.4% 提升至 68.2%)和 OpenCode(57.2% 提升至 66.6%)上的性能提升,同时保持 rollout-训练概率相关性高于 0.99。

英文摘要:

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train-inference discrepancies decouple rollout behavior from policy updates. To address this, we present LEGO-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow. LEGO-RL is built upon three pillars: (1) faithful optimization via in-process LLM proxying that captures raw generation streams for token-level alignment and robust trainer-side log-probability recomputation, even under harness-side compaction or re-serialization; (2) reliable execution via scalable sandbox orchestration featuring image caching and stage-wise defenses to mitigate reward hacking; and (3) observable training through an integrated plugin that automates validation and monitoring, paired with a Live UI for granular trajectory diagnostics. We evaluate LEGO-RL by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses. LEGO-RL improves Qwen3.5-35B-A3B across OpenHands SDK (64.0% to 70.4%), Claude Code (62.4% to 68.2%), and OpenCode (57.2% to 66.6%) on SWE-bench Verified, while maintaining a rollout-training probability correlation above 0.99.

补充信息

相关深度报道

↑