arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可扩展的上下文强化学习:基于循环算法蒸馏

Scalable In-Context Reinforcement Learning with Recurrent Algorithm Distillation

Yuanqing Ma, Zhenrui Zheng, Chenjun Xiao

arXiv 2609.35333首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出循环算法蒸馏(RAD),通过压缩历史为固定大小潜在记忆,在显著减小上下文窗口的同时达到与标准AD相当的性能,实现可扩展的上下文强化学习。

AI 中文摘要

算法蒸馏(AD)已展示了Transformer在不进行显式权重更新的情况下执行上下文强化学习的卓越能力。然而,捕捉长期学习进展需要庞大的上下文窗口,这带来了高昂的内存成本,并限制了在复杂、长时程任务中的可扩展性。为解决这一瓶颈,我们提出了循环算法蒸馏(RAD)。RAD采用双组件架构:一个压缩Transformer,将扩展的交互历史蒸馏为紧凑的潜在令牌;以及一个AD Transformer,利用这些压缩记忆与近期转换的混合上下文自回归地生成动作。通过维持固定大小的潜在缓冲区,RAD将有效历史长度与计算复杂度解耦,在功能上为模型提供了长时程记忆。跨多种环境的实证评估表明,RAD在显著减小上下文窗口尺寸的同时,达到了与标准AD相当的性能,为高效的上下文决策提供了一种可扩展的解决方案。

英文摘要

Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scalability in complex, long-horizon tasks. To address this bottleneck, we propose Recurrent Algorithm Distillation (RAD). RAD employs a dual-component architecture: a Compression Transformer that distills extended interaction histories into compact latent tokens, and an AD Transformer that auto-regressively generates actions using a hybrid context of these compressed memories and recent transitions. By maintaining a fixed-size latent buffer, RAD decouples the effective history length from computational complexity, functionally providing the model with a long-horizon memory. Empirical evaluations across diverse environments demonstrate that RAD matches the asymptotic performance of standard AD with significantly reduced context window sizes, offering a scalable solution for efficient in-context decision-making.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑