arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34488cs.LGcs.AI

FlexLoop: 深度弹性循环策略用于深度强化学习中的自适应测试时计算

FlexLoop: Depth-Elastic Looped Policies for Adaptive Test-Time Computation in Deep RL

Xun Wang, Ruishuo Chen, Yu Chen, Zhuoran Li, Longbo Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对预训练循环策略在深度上特化导致无法弹性推理的问题,提出FlexLoop训练后框架,通过相邻深度蒸馏实现深度弹性,在保持性能的同时平均深度降低43%,速度提升1.34倍。

中文摘要 AI 辅助

循环架构通过在不同循环步骤中重用相同参数来扩展计算量,近期研究表明,这类架构在长视界任务上显著提升了深度强化学习策略的性能。由于循环深度直接控制计算量,人们可能预期循环策略能够自然支持跨循环深度的弹性推理。然而,令人惊讶的是,我们发现预训练的循环策略表现出严重的循环深度特化现象:可靠的决策集中在接近完整训练深度的区域,即使较少的计算量可能足够,部署时的计算量也被束缚在该深度上。因此,实现深度弹性(即在部署时通过自适应计算在不同循环深度下做出可靠决策)仍然是一个关键挑战。为解决这一问题,我们提出了FlexLoop,一种新颖的训练后框架,将预训练的固定深度循环策略转换为深度弹性策略。FlexLoop在保持原始强化学习目标训练以保留全深度能力的同时,执行相邻深度策略蒸馏,逐步将决策质量从较深循环步骤转移到较浅循环步骤。所得策略支持跨循环深度的可靠推理,并通过循环深度一致性实现状态级自适应推理。在30个在线和离线长视界目标条件环境中进行的实验表明,FlexLoop在保持全深度性能的同时,使较浅深度变得有效。在保持竞争性性能的情况下,FlexLoop将平均循环深度最多降低43%,并在压力测试中实现高达1.34倍的墙钟加速。

英文摘要

Looped architectures scale computation by reusing the same parameters across recurrent steps, and recent work shows that they substantially improve deep reinforcement learning policies on long-horizon tasks. Since recurrent depth directly controls computation, one may expect looped policies to naturally support elastic inference across recurrent depths. Surprisingly, we find that pretrained looped policies exhibit severe recurrent-depth specialization: reliable decisions are concentrated near the full trained depth, tying deployment computation to this depth even when less computation may suffice. Achieving depth elasticity, i.e., reliable decisions across recurrent depths with adaptive computation at deployment, therefore remains a key challenge. To address this, we propose FlexLoop, a novel post-training framework that converts pretrained fixed-depth looped policies into depth-elastic policies. FlexLoop keeps training on the original RL objective to preserve full-depth capability while performing adjacent-depth policy distillation to progressively transfer decision quality from deeper to shallower recurrent steps. The resulting policy supports reliable inference across recurrent depths and enables state-wise adaptive inference through recurrent-depth consistency. Experiments on $30$ online and offline long-horizon goal-conditioned environments show that FlexLoop preserves full-depth performance while making shallower depths effective. Keeping competitive performance, FlexLoop reduces average recurrent depth by up to $\bf{43\%}$ and achieves up to $\bf{1.34\times}$ wall-clock speedup in a stress test.

发表机构

  • Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑