arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于复杂序列决策任务的混合大语言模型增强强化学习智能体

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks

Christophe D. Hounwanou, John Emeka Eze, Yaé Ulrich Gaba

arXiv 2608.03502首次发表:更新:

AI 中文总结

本文提出混合LLM增强RL智能体,融合LLM规划与RL动作优化,经实验验证其在序列决策任务上优于仅RL、仅LLM的基线方法,为构建更强自主系统提供了新方向。

AI 中文摘要

大语言模型(LLMs)近期展现出强大的推理、规划和工具使用能力,催生了新型自主智能体,但基于LLM的智能体在需要精准动作优化与环境交互的长程序列决策任务中表现不佳;强化学习(RL)虽对序列控制有效,却缺乏复杂场景所需的高层抽象与任务分解能力。本文提出一种LLM增强强化学习智能体,将LLM驱动的规划与RL驱动的动作优化相融合,该架构利用LLM生成子目标、结构化规划与上下文指导,同时RL智能体通过与环境交互优化底层动作。在序列决策任务上的实验表明,与仅RL和仅LLM的基线方法相比,该方法样本效率更高、成功率更高、动作轨迹更连贯,这种混合范式为构建更强大的自主系统指明了有前景的方向。

英文摘要

Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents. However, LLM-based agents struggle with long-horizon sequential decision tasks that require precise action optimization and environment interaction. Reinforcement Learning (RL), while effective for sequential control, often lacks the high-level abstraction and task decomposition abilities needed for complex scenarios. This paper introduces an LLM-Augmented Reinforcement Learning Agent that integrates LLM-driven planning with RL-based action optimization. The proposed architecture leverages the LLM to generate subgoals, structured plans, and contextual guidance, while the RL agent refines low-level actions through interaction with the environment. Experiments on sequential decision tasks demonstrate improved sample efficiency, higher success rates, and more coherent action trajectories compared to RL-only and LLM-only baselines. This hybrid paradigm highlights a promising direction for building more capable autonomous systems.

CommentsThis submission is withdrawn because the uploaded manuscript does not accurately reflect the intended structure or results. Several components referenced in the text are incomplete or not represented in the PDF, and the current version may mislead readers. The work is therefore withdrawn to maintain clarity of the record

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑