arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LEACL:用于长时程操作任务强化学习的大语言模型增强自动课程学习

LEACL: LLM-Enhanced Automatic Curriculum Learning for Reinforcement Learning in Long-Horizon Manipulation Tasks

Faraz Heravi, James Ouyang, Zifan Xu, Arjun Kumar, Yoonchang Sung, Peter Stone

arXiv 2607.23515首次发表:更新:

发表机构

The University of Texas at Austin; Nanyang Technological University; Sony AI(德克萨斯大学奥斯汀分校; 南洋理工大学; 索尼人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程操作任务强化学习挑战,提出LEACL框架,结合大语言模型与自动课程学习,用大语言模型分解任务并生成规范,由自动课程学习算法仅依稀疏奖励信号指导学习,在相关任务上取得更好渐近性能。

AI 中文摘要

长时程操作任务因稀疏奖励信号和长时程对强化学习构成重大挑战。自动课程学习(ACL)通过按顺序从易到难训练智能体来应对,但依赖人工制定的任务相关规范。大语言模型(LLM)的进展为任务分解提供了可能,但现有基于LLM的方法依赖手工设计的密集奖励函数。本文提出LLM增强自动课程学习(LEACL)框架,用LLM分解任务并生成规范,由ACL算法仅用稀疏奖励信号指导学习。在LIBERO基准的五个长时程操作任务上评估,LEACL在成功率方面取得更好的渐近性能。

英文摘要

Long-horizon manipulation tasks pose significant challenges for reinforcement learning due to sparse reward signals and long horizons. Automatic curriculum learning (ACL) has been proposed to tackle these challenges by progressively training agents on a sequence of tasks, from easier to more difficult. However, the success of ACL depends heavily on task-dependent specifications-such as well-defined task parameter spaces and difficulty measures-which are often manually crafted and difficult to generalize across diverse tasks. Recent advances in large language models (LLMs) offer a promising alternative by enabling the decomposition of complex tasks into meaningful subtasks using the LLMs' web-scale common-sense knowledge. This decomposition can provide a natural curriculum structure for efficient learning of long-horizon tasks. However, existing LLM-based methods typically rely on hand-designed dense reward functions to learn each subtask, which can introduce bias and still requires significant human supervision. In this work, we propose LLM-enhanced automatic curriculum learning (LEACL), a framework that integrates LLMs and ACL to address these limitations. Specifically, LLMs are used to both decompose tasks into subtasks and to generate task-dependent specifications for each subtask. These specifications are then used by ACL algorithms to guide learning using only sparse reward signals, eliminating the need for dense reward design. We evaluate LEACL on five long-horizon manipulation tasks from the LIBERO benchmark. LEACL achieves better asymptotic performance in terms of the success rates compared to human-designed dense rewards.

Comments8 pages, 4 figures, Published in the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑