arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2503.19007cs.LGcs.AI

使用LLM引导的语义分层强化学习进行选项发现

Option Discovery Using LLM-guided Semantic Hierarchical Reinforcement Learning

  • University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

Chak Lam Shek, Pratap Tokekar

更新

AI总结:

本文提出LDSC框架,利用LLM引导的子目标选择和选项重用,通过三阶段分层强化学习提升复杂机器人任务的样本效率、泛化能力和多任务适应性,平均奖励比基线高55.9%。

AI中文摘要:

大型语言模型(LLMs)在推理和决策方面展现出了显著的前景,但它们与强化学习(RL)的集成在复杂机器人任务中仍未得到充分探索。在本文中,我们提出了一种LLM引导的分层强化学习框架,称为LDSC,该框架利用LLM驱动的子目标选择和选项重用,以提高样本效率、泛化能力和多任务适应性。传统的RL方法常常遭受探索效率低下和计算成本高昂的问题。分层RL有助于解决这些挑战,但现有方法在面对新任务时往往无法有效地重用选项。为了解决这些局限性,我们引入了一个三阶段框架,该框架利用LLM根据任务的自然语言描述生成子目标,一种可重用的选项学习和选择方法,以及一个动作级策略,从而在多样化的任务中实现更有效的决策。通过将LLM纳入子目标预测和策略指导,我们的方法提高了探索效率并增强了学习性能。平均而言,LDSC在平均奖励方面比基线高出55.9%,证明了其在复杂RL设置中的有效性。更多细节和实验视频可在以下链接中找到:https://raaslab.org/projects/LDSC。

英文摘要:

Large Language Models (LLMs) have shown remarkable promise in reasoning and decision-making, yet their integration with Reinforcement Learning (RL) for complex robotic tasks remains underexplored. In this paper, we propose an LLM-guided hierarchical RL framework, termed LDSC, that leverages LLM-driven subgoal selection and option reuse to enhance sample efficiency, generalization, and multi-task adaptability. Traditional RL methods often suffer from inefficient exploration and high computational cost. Hierarchical RL helps with these challenges, but existing methods often fail to reuse options effectively when faced with new tasks. To address these limitations, we introduce a three-stage framework that uses LLMs for subgoal generation given natural language description of the task, a reusable option learning and selection method, and an action-level policy, enabling more effective decision-making across diverse tasks. By incorporating LLMs for subgoal prediction and policy guidance, our approach improves exploration efficiency and enhances learning performance. On average, LDSC outperforms the baseline by 55.9\% in average reward, demonstrating its effectiveness in complex RL settings. More details and experiment videos could be found in \href{https://raaslab.org/projects/LDSC/}{this link\footnote{https://raaslab.org/projects/LDSC}}.

↑