发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出RLTL;DR,通过让智能体在失败后自我生成TL;DR见解并内化任务到见解的映射,突破自我改进中低成功率的学习瓶颈,在困难工具调用和编码任务上显著提升Pass@1,并展示了紧凑训练范式SFTL;DR的潜力。
AI 中文摘要
带可验证奖励的强化学习(RLVR)的常见范式是让智能体对任务进行多次尝试,并针对成功的尝试进行优化。这在自我改进领域变得有问题,因为任务难度极高,智能体成功的机会很低甚至为零,而且没有教师模型或示例解决方案可供提炼。在本文中,我们引入了RLTL;DR。在每次失败尝试后,我们向策略展示验证器的输出,并让其以单条TL;DR见解的形式写出自己的反馈。下一次采样(rollout)以之前所有见解为条件,我们顺序采样直到找到解决方案。此外,我们支持对上下文中的见解进行反向传播,以内化直接的任务到见解的映射。在具有挑战性的工具调用和编码数据集(过滤到Pass@128=0)上,对Qwen 3.5 9B Thinking策略进行标准GRPO训练,其Pass@1保持在0%到1%的平坦水平。RLTL;DR突破了这一学习障碍,在训练时上下文中有见解的情况下实现了14-31%的Pass@1,并且关键的是,在评估时上下文中没有见解的情况下实现了12-13%的Pass@1。我们确定关键是任务到见解的内化。为了进一步研究这一点,我们将我们的方法简化为SFTL;DR,仅训练(任务,见解)对,而不展示或反向传播任何采样。仅用4k个这样的对进行训练,几乎恢复了RLTL;DR和经典SFT在完整采样上的全部性能。这展示了一种有前景的紧凑训练范式,形式为“在这种任务上,记住这类事情”,我们希望这能激发未来的研究。
英文摘要
The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form "on this sort of task, keep this sort of thing in mind", which we hope to inspire future research on.