循环Actor:用于强化学习的深度递归推理模型
Looped Actor: Depth-Recurrent Reasoning Models for Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本研究提出循环Actor,一种基于Transformer的策略,通过动态循环共享计算块实现自适应计算,在22个任务上匹配或超越16倍参数基线,显著提升强化学习智能体的规划能力。
中文摘要 AI 辅助
循环推理模型重复应用一组共享参数,从而在不增加模型大小的情况下实现更多计算。这些模型还通过动态决定何时停止循环来支持依赖于输入的计算。受循环Transformer在语言建模和推理中近期成功的启发,我们研究动态循环是否同样能有益于顺序决策。我们通过证明存在马尔可夫决策过程,其中状态自适应策略以渐近更少的期望计算达到最优回报,而任何最优固定运行时间策略则不能,从而为这种方法提供了复杂性理论动机。为了在实践中学习计算自适应策略,我们引入了循环Actor,一种基于Transformer的策略,它使用共享计算块将潜在表示反复精炼至固定点。这允许模型通过根据当前状态改变循环次数来自适应地分配计算。我们在六个环境中的22个任务上评估循环Actor,范围从组合谜题到机器人操作,并涵盖在线和离线强化学习(RL),包括离散和连续动作。循环Actor匹配或超过了具有16倍参数的非共享基线性能,在动作选择需要大量多步规划的环境中增益最大。对于Boxoban环境,我们发现计算分配是有结构的:循环次数随着剩余推数的增加而增加,并且未来的最优推数在连续循环中从潜在状态变得越来越可预测。总之,这些结果突显了循环Actor作为一种简单高效的方式,为RL智能体配备自适应计算并提升其规划能力。代码可在以下https URL获取。
英文摘要
Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also support input-dependent computation by dynamically deciding when to stop looping. Motivated by the recent success of looped transformers in language modeling and reasoning, we investigate whether dynamic looping can similarly benefit sequential decision-making. We provide a complexity-theoretic motivation for this approach by showing that there exist Markov decision processes in which a state-adaptive policy achieves the optimal return with asymptotically less expected computation than any optimal fixed-runtime policy. To learn compute-adaptive policies in practice, we introduce Looped Actor, a transformer-based policy that repeatedly refines a latent representation toward a fixed point using a shared computational block. This allows the model to allocate computation adaptively by varying the number of loops based on the current state. We evaluate Looped Actor on 22 tasks across six environments, ranging from combinatorial puzzles to robotic manipulation and spanning online and offline reinforcement learning (RL) with discrete and continuous actions. Looped Actor matches or exceeds the performance of an untied baseline with 16$\times$ more parameters, with the largest gains in environments where action selection requires substantial multistep planning. For the Boxoban environment, we find that the computation allocation is structured: the number of loops increases with the number of remaining pushes and future optimal pushes become increasingly predictable from the latent state over successive loops. Together, these results highlight actor looping as a simple and efficient way to equip RL agents with adaptive computation and improve their planning capabilities. Code is available at https://github.com/camail-official/LoopedActor
发表机构
- ELLIS Institute Tübingen(ELLIS研究所蒂宾根)
- Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
- Tübingen AI Center(蒂宾根人工智能中心)
- Liquid AI
- CSAIL, MIT(麻省理工学院计算机科学与人工智能实验室)
- Case Western Reserve University(凯斯西储大学)
机构由 AI 辅助整理,请以论文原文为准。