AI 中文总结
本文提出CANOPY协议,通过解决信号匮乏与策略漂移问题,使仅基于结果的强化学习在小型开放模型上实现长视距交互任务,在AppWorld和SWE-bench Verified基准上取得优异性能。
AI 中文摘要
强化学习是对大语言模型(LLM)智能体进行后训练以完成仅由任务结束时验证评判的长视距交互任务的自然方式,但普遍观点认为,仅基于结果的强化学习在小型开放模型上很快会达到性能上限。因此,近期研究通过使用更密集的奖励、监督微调(SFT)先验、技能库、精心设计的记忆或多智能体编排等方式对训练进行补充。本文认为该上限是常规实践中两种缺陷导致的假象:信号匮乏——采用稀疏仅基于结果奖励的组相对强化学习仅在任务的回滚组混合成功与失败时才会产生梯度,因此规模不足的探索会恰好抑制最难、最具指导性的任务;策略漂移——从小型任务池中挤出过多更新会导致策略本身退化,因为当饱和使有信息的组变得罕见时,无锚定的目标会让采样分布恰好崩溃。本文提出CANOPY(Coverage-ANchored On-PolicY RL,覆盖锚定的同策略强化学习),这是一种直接针对上述两种缺陷的极简协议:将相同任务的探索规模扩大,直到自然信号重新出现,保持每次更新同策略、带有KL散度锚定且仅局限于智能体自身的动作标记,随后在测试时利用扩大的交互预算。在长视距交互编码基准AppWorld上,仅通过环境交互(无需任务特定监督、辅助信用信号或复杂智能体支架)使用CANOPY训练的Qwen3-14B策略登上了公共排行榜(2026年2月;Test-Normal TGC为86.9,Test-Challenge为67.6),相同设计原则使Qwen3.5-9B在SWE-bench Verified上的得分提升了16.6个百分点。因此,仅智能体强化学习就能将长视距能力直接内化到小型开放模型中;本文计划在该https://URL发布完整的训练栈。
英文摘要
Reinforcement learning is a natural way to post-train LLM agents for long-horizon interactive tasks judged only by end-of-task verification, yet a shared belief holds that outcome-only RL soon hits a ceiling on small open models. Recent work therefore compensates around the training with denser rewards, SFT priors, skill libraries, curated memory, or multi-agent orchestration. We argue the ceiling is an artifact of two failures of common practice. Signal starvation: group-relative RL with sparse outcome-only rewards yields a gradient only when a task's rollout group mixes successes and failures, so under-scaled exploration silences exactly the hardest, most instructive tasks. Policy drift: squeezing many updates out of a small task pool degrades the policy itself, as an unanchored objective lets the sampling distribution collapse exactly when saturation has already made informative groups rare. We present CANOPY (Coverage-ANchored On-PolicY RL), a minimalist protocol attacking both directly: scale same-task exploration until the natural signal reappears, keep every update on-policy, KL-anchored, and confined to the agent's own action tokens, then cash in an enlarged interaction budget at test time. On AppWorld, a long-horizon interactive coding benchmark, a Qwen3-14B policy trained with CANOPY through environment interaction alone--without task-specific supervision, auxiliary credit signals, or elaborate agent scaffolding--topped the public leaderboard (Feb. 2026; Test-Normal TGC 86.9, Test-Challenge 67.6), and the same design principles lift Qwen3.5-9B on SWE-bench Verified by 16.6 points. Agentic RL alone thus internalizes long-horizon capability directly into a small open model; we plan to release the complete training stack at https://github.com/AlibabaResearch/SignalCoverageRL.
Comments13 pages, 6 figures