Agent Lightning v1.0:走向可控的智能体强化学习
Agent Lightning v1.0: Towards Harnessed Agentic RL
浏览论文内容
中文总结 AI 辅助
研究针对可控智能体强化学习的挑战,推出轻量级框架 Agent Lightning v1.0,经评估可将 Qwen3.5-9B 在 SWE-bench Verified 上的性能提升14.6个百分点,发布完整流程脚本以推进相关可复现研究。
中文摘要 AI 辅助
现代智能体在管理工具、上下文和控制流的智能体 harness(管控框架)中运行,这使得管控框架成为智能体系统的关键组成部分。最初的 Agent Lightning 引入了一种解耦架构,通过大语言模型(LLM)端点代理将任意智能体连接到强化学习(RL)训练,该方法后来被 verl、Uni-Agent、AReaL 2.0、slime 和 Polar 等框架采用。我们将这种范式称为可控智能体强化学习,其中部署阶段的管控框架直接参与模型的后训练。可控智能体强化学习与传统智能体强化学习存在根本差异:管控框架而非训练引擎掌控环境交互循环,训练器仅观测 LLM 的请求-响应对序列。这在重新分词、样本合并、优势计算、损失归一化和后端调度方面带来了挑战,可能显著影响训练的稳定性和有效性。我们推出 Agent Lightning v1.0,这是一个用于可控智能体强化学习的轻量级框架,代码量约为 3500 行。它支持任意智能体管控框架,并作为研究上述挑战的实用测试平台。我们在遵循指令、搜索和编码智能体上对其进行评估,并提供了用于编码智能体强化学习的完整可复现流程。仅使用 6000 个训练样本和适度的计算资源,RL 将 Qwen3.5-9B 在 SWE-bench Verified 上的性能从 41.8% 提升至 56.4%,绝对提升了 14.6 个百分点。我们发布了完整的工作流和训练脚本,以促进可控智能体强化学习的可复现研究。
英文摘要
Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through an LLM endpoint proxy, an approach later adopted by frameworks such as verl Uni-Agent, AReaL 2.0, slime, and Polar. We refer to this paradigm as harnessed agentic RL, where the deploy-time harness directly participates in model post-training. Harnessed agentic RL differs fundamentally from traditional agentic RL: the harness, rather than the training engine, owns the environment interaction loop, while the trainer observes only sequences of LLM request-response pairs. This introduces challenges in retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling, which can substantially affect training stability and effectiveness. We present Agent Lightning v1.0, a lightweight framework for harnessed agentic RL implemented in approximately 3,500 lines of code. It supports arbitrary agent harnesses and serves as a practical testbed for studying these challenges. We evaluate it on instruction-following, search, and coding agents, and provide a complete reproducible pipeline for coding-agent RL. Using only 6K training examples and modest compute, RL improves Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%, a 14.6-point absolute gain. We release the complete workflow and training scripts to facilitate reproducible research on harnessed agentic RL.
发表机构
- Microsoft(微软)
- Fudan University(复旦大学)
- Zhejiang University(浙江大学)
- University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。