arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17528cs.AIcs.SE

Agent Lightning v1.0:走向可控的智能体强化学习

Agent Lightning v1.0: Towards Harnessed Agentic RL

Zhiyuan He, Siwei Zhang, Zhiwen Zhou, Yuqing Yang, Yu Kang, Yuge Zhang, Luna K. Qiu, Tin Yan Tsui, Jiahang Xu, Chong Luo

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对可控智能体强化学习的挑战,推出轻量级框架 Agent Lightning v1.0,经评估可将 Qwen3.5-9B 在 SWE-bench Verified 上的性能提升14.6个百分点,发布完整流程脚本以推进相关可复现研究。

中文摘要 AI 辅助

现代智能体在管理工具、上下文和控制流的智能体 harness(管控框架)中运行,这使得管控框架成为智能体系统的关键组成部分。最初的 Agent Lightning 引入了一种解耦架构,通过大语言模型(LLM)端点代理将任意智能体连接到强化学习(RL)训练,该方法后来被 verl、Uni-Agent、AReaL 2.0、slime 和 Polar 等框架采用。我们将这种范式称为可控智能体强化学习,其中部署阶段的管控框架直接参与模型的后训练。可控智能体强化学习与传统智能体强化学习存在根本差异:管控框架而非训练引擎掌控环境交互循环,训练器仅观测 LLM 的请求-响应对序列。这在重新分词、样本合并、优势计算、损失归一化和后端调度方面带来了挑战,可能显著影响训练的稳定性和有效性。我们推出 Agent Lightning v1.0,这是一个用于可控智能体强化学习的轻量级框架,代码量约为 3500 行。它支持任意智能体管控框架,并作为研究上述挑战的实用测试平台。我们在遵循指令、搜索和编码智能体上对其进行评估,并提供了用于编码智能体强化学习的完整可复现流程。仅使用 6000 个训练样本和适度的计算资源,RL 将 Qwen3.5-9B 在 SWE-bench Verified 上的性能从 41.8% 提升至 56.4%,绝对提升了 14.6 个百分点。我们发布了完整的工作流和训练脚本,以促进可控智能体强化学习的可复现研究。

英文摘要

Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through an LLM endpoint proxy, an approach later adopted by frameworks such as verl Uni-Agent, AReaL 2.0, slime, and Polar. We refer to this paradigm as harnessed agentic RL, where the deploy-time harness directly participates in model post-training. Harnessed agentic RL differs fundamentally from traditional agentic RL: the harness, rather than the training engine, owns the environment interaction loop, while the trainer observes only sequences of LLM request-response pairs. This introduces challenges in retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling, which can substantially affect training stability and effectiveness. We present Agent Lightning v1.0, a lightweight framework for harnessed agentic RL implemented in approximately 3,500 lines of code. It supports arbitrary agent harnesses and serves as a practical testbed for studying these challenges. We evaluate it on instruction-following, search, and coding agents, and provide a complete reproducible pipeline for coding-agent RL. Using only 6K training examples and modest compute, RL improves Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%, a 14.6-point absolute gain. We release the complete workflow and training scripts to facilitate reproducible research on harnessed agentic RL.

发表机构

  • Microsoft(微软)
  • Fudan University(复旦大学)
  • Zhejiang University(浙江大学)
  • University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑