发表机构
Hansung University(汉城大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长程推理中智能体因中间错误偏离目标的问题,提出SGG-ReflAct,将单路径或束搜索LLM规划器生成的子目标融入反思过程,在ALFWorld和ScienceWorld上分别提升14.9和8.0个百分点成功率,并减少幻觉动作。
AI 中文摘要
近期推理骨干网络的进展使大语言模型(LLM)智能体能够处理复杂的多步任务。然而,随着推理范围的增长,不一致的内部信念会导致中间错误,使智能体偏离其目标。这一局限在REFLACT中同样存在,该方法在每一步仅反思最终目标,而未显式考虑中间子目标。为解决此问题,我们提出SGG-ReflAct(子目标引导反思),一种将单路径LLM规划器生成的子目标整合进反思过程的推理骨干网络。我们进一步将该框架扩展为BeamSGG-ReflAct,用基于束搜索的LLM规划器替代单路径规划器,以进行结构化计划探索。我们在ALFWorld、ScienceWorld和Jericho上使用多个LLM模型进行了实验。SGG-ReflAct在几乎所有设置中均优于REFLACT,在ALFWorld上使用Llama-3.1-8B-Instruct取得了最高14.9个百分点的成功率提升,在ScienceWorld上取得了8.0个百分点的提升。我们的实验分析表明,SGG-ReflAct减少了幻觉动作,并在程序化有序任务上获得了最大收益。此外,BeamSGG-ReflAct的实验结果表明,该骨干网络的有效性取决于计划质量:显式指定所需操作可恢复仅靠计划搜索无法获得的收益。这些结果证明,SGG-ReflAct提供了一种实用且高效的推理骨干网络,通过易于集成的方式使LLM智能体在复杂长程任务中实现可靠性能。
英文摘要
Recent advances in reasoning backbones have empowered large language model (LLM)agentstotackle complex, multi-step tasks. However, as reasoning horizons grow, inconsistent internal beliefs induce intermediate errors that cause agents to drift from their goals. This limitation also persists in REFLACT, which reflects only on the end-goal at each step without explicitly considering intermediate sub goals. To address this problem, we propose SGG-ReflAct (Sub-Goal Guided Re flAct), a reasoning backbone that integrates sub-goals generated through a single path LLM planner into the reflection process. We further extend this framework to BeamSGG-ReflAct, which replaces the single-path planner with a beam search based LLM planner for structured plan exploration. We run experiments on ALF World, ScienceWorld, and Jericho with multiple LLM models. SGG-ReflAct out performs REFLACT in nearly all settings, achieving best success rate gains of 14.9 percentage points on ALFWorld and 8.0 percentage points on ScienceWorld with Llama-3.1-8B-Instruct. Our experimental analysis shows that SGG-ReflAct re duces hallucinated actions and achieves its largest gains on procedurally ordered tasks. Furthermore, experimental results with BeamSGG-ReflAct show that the backbone's effectiveness depends on plan quality: explicitly specifying the re quired operations recovers gains that plan searching alone cannot achieve. These results demonstrate that SGG-ReflAct offers a practical and highly effective rea soning backbone, enabling LLM agents to achieve reliable performance in com plex, long-horizon tasks through easy integration.