arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从轨迹到指令:语言条件元强化学习

From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning

Garvit Singla, Uma Maheswari Natarajan, Raghuram Bharadwaj Diddigi

arXiv 2607.18830首次发表:更新:

发表机构

International Institute of Information Technology Bangalore(班加罗尔国际信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何改进强化学习中模型无关元学习的内循环适应机制,提出LA - MAML,利用任务语言指令替代传统的轨迹收集和梯度更新,实验证明其在降低训练时间的同时性能相当或更优。

AI 中文摘要

模型无关元学习(MAML)是强化学习中广泛使用的框架,通过学习可快速适应新任务的全局策略参数实现高效迁移。MAML训练分两个循环进行,传统上内循环通过从任务环境收集轨迹并对经验期望回报应用梯度更新来执行,成本较高。我们注意到是外循环驱动全局参数的实际学习,所以内循环适应机制不必限于基于梯度的。基于此,我们提出LA - MAML,利用任务语言指令作为特定于任务的信号,通过学习的任务指令嵌入一步调整全局策略参数来修改内循环,取代轨迹收集和基于梯度的更新。在BabyAI基准测试上的实验表明,LA - MAML在每次迭代的挂钟训练时间显著更低的情况下,性能与基线相当或有所提高。这些结果表明语言指令是元强化学习中基于轨迹的内循环适应的有效且高效替代。

英文摘要

Model-Agnostic Meta-Learning (MAML) is a widely used framework for reinforcement learning (RL) that enables efficient transfer by learning global policy parameters that can be rapidly adapted to new tasks. MAML training proceeds in two loops: an inner loop where the global parameters are adapted to task-specific parameters, and an outer loop where these task-specific parameters are evaluated and losses are back-propagated to improve the global parameters. Traditionally, the inner loop adaptation is performed by collecting trajectories from the task environment and applying gradient updates on the empirical expected return, which can be a costly operation. We note that it is the outer loop that drives the actual learning of global parameters, and therefore the inner loop adaptation mechanism need not be restricted to be gradient-based. This observation leads us to ask: Can we replace the inner loop trajectory collection and gradient update with a simpler, task-specific signal? In many practical settings, tasks are naturally accompanied by language instructions. Leveraging these instructions as a direct task-specific signal, we propose LA-MAML (Language Adapted MAML), which modifies the inner loop by adapting the global policy parameters in a single step through a learned embedding of the task instruction, replacing the inner loop trajectory collection and gradient-based updates. Experiments on the BabyAI benchmark demonstrate that LA-MAML achieves competitive or improved performance compared to baselines at a significantly lower per-iteration wall-clock training time. These results demonstrate that language instructions are an effective and efficient substitute for trajectory-based inner loop adaptation in meta RL.

CommentsAccepted at the International Conference on Artificial Neural Networks (ICANN 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑