KV-streams:用于智能体强化学习中高效压缩的键值流
KV-streams for Efficient Compaction in Agentic Reinforcement Learning
- Mila(米拉研究所)
- Microsoft(微软)
- McGill University(麦吉尔大学)
- Université de Montréal(蒙特利尔大学)
- Polytechnique Montréal(蒙特利尔理工学院)
- ServiceNow Inc(ServiceNow公司)
- Rensselaer Polytechnic Institute(伦斯勒理工学院)
- University College London, University of London(伦敦大学学院)
- Vmax(Vmax公司)
- Edinburgh University(爱丁堡大学)
- Microsoft Research(微软研究院)
- Cohere(Cohere公司)
- HEC Montréal(蒙特利尔高等商学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对智能体强化学习中上下文压缩导致的训练吞吐瓶颈,提出KV-streams策略,通过前向流式传输KV缓存替代清空,实现2.6至5倍训练加速且不损性能,并发现流式缓存可充当循环状态传递长期信息。
AI中文摘要:
扩展智能体大语言模型(LLM)的交互范围受到GPU内存中需容纳越来越长上下文轨迹的限制。上下文压缩一直是缓解此问题的最常用机制,能在给定轨迹下保持GPU内存恒定。然而,大多数压缩策略依赖于多次对LLM上下文进行预填充,这阻碍了训练吞吐量。为缓解此瓶颈并实现高效的、可训练的压缩,我们提出了KV-streams,这是一种即插即用的策略,与任何压缩策略兼容,能显著提高吞吐量,且未显示任何阻碍性能的证据。KV-streams通过前向传输KV缓存而非在每次压缩后将其清空,从而实现可扩展的压缩。我们展示了KV-streams支持三种不同的压缩策略,在训练中实现了2.6至5倍的挂钟加速。除了效率提升,我们还发现流式KV缓存可充当循环状态,携带早已从上下文中消失的信息。具体而言,在受控环境中,我们表明与先前工作相反,仅凭强化学习(RL)就足以使该行为出现。总体而言,我们证明KV-streams是任何后训练流程中一种高效、轻量级的即插即用补充。
英文摘要:
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.