AI 中文总结
研究针对LLM智能体应用挑战,设计优化工作流程,合成多种AI范式,引入POMDP路由和自我修正奖励模型,整合多模态输入与强化学习原则,实验显示相比主流基线任务成功率等提升24.5%,为自主系统AI技术开发提供参考框架。
AI 中文摘要
本文通过设计和优化智能体工作流程,解决当前大语言模型(LLM)智能体应用中的关键技术挑战,包括长期规划、稀疏奖励归因和动态环境交互。所提出的架构基于核心人工智能范式的合成。与传统依赖静态提示且缺乏强大感知 - 行动循环的基线模型不同,引入了部分可观测马尔可夫决策过程(POMDP)路由机制,并辅以内部自我修正奖励模型。通过整合多模态输入和先进强化学习原则,智能体保持长期结构记忆并动态调整推理路径。在ALFWorld实体模拟环境和WebShop在线导航基准上的实证实验表明,任务成功率和轨迹效率比标准ReAct框架等主流基线有24.5%的绝对提升。综合消融研究证实奖励驱动批判模块在抑制幻觉率方面的重要贡献。该研究将强化学习和基于图的记忆的理论基础与自主智能体工作流程相连接,为复杂多步自主系统中的人工智能技术开发提供了实用、可扩展的参考框架。代码可在指定网址获取。
英文摘要
This paper addresses key technical challenges in current large language model (LLM) agent applications, including long-horizon planning, sparse reward attribution, and dynamic environmental interaction, by designing and optimizing an intelligent agent workflow. The proposed architecture is based on the synthesis of core AI paradigms: Visual, Language, Generative, Graph, Multimodal, Reinforcement, and Agent Intelligence. Unlike conventional baseline models that rely on static prompting and lack robust perception-action loops, our approach introduces a Partially Observable Markov Decision Process (POMDP) routing mechanism. This mechanism is augmented with an internal, self-correcting reward model that evaluates decision trajectories before execution. By integrating multimodal inputs and advanced reinforcement learning principles (such as proximal policy optimization and value function approximation), the agent maintains long-term structural memory and dynamically adapts its reasoning pathways to mitigate error accumulation. Empirical experiments on the ALFWorld embodied simulation environment and the WebShop online navigation benchmark demonstrate a 24.5% absolute improvement in task success rate and trajectory efficiency over mainstream baselines like the standard ReAct framework. Comprehensive ablation studies confirm the significant contribution of the reward-driven critique module in suppressing hallucination rates. This research bridges theoretical foundations of reinforcement learning and graph-based memory with autonomous agent workflows. Ultimately, the resulting architecture offers a practical, scalable reference framework for developing artificial intelligence technologies in complex, multi-step autonomous systems. Code is available at https://github.com/01Amez/RLAW_Implementation.
Comments29 pages, 1 table, native TikZ/pgfplots diagrams. Code available at https://github.com/01Amez/RLAW_Implementation