arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00455cs.AI

面向大语言模型智能体的基于信念的世界模型

Towards a Belief-Based World Model for LLM Agents

发表机构伊利诺伊大学厄巴纳-香槟分校 · 国际商业机器公司
查看机构详情
  • UIUC(伊利诺伊大学厄巴纳-香槟分校)
  • IBM(国际商业机器公司)

机构由 AI 辅助整理,请以论文原文为准。

Shubham Kumar, Harshit Kumar, Narendra Ahuja, Saurabh Jha

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对LLM智能体在部分可观测下的决策缺陷,提出BB-WMs,将世界模型的信念暴露给LLM策略,可提升部分可观测任务性能,且与现有基于模拟的世界模型互补。

中文摘要 AI 辅助

大语言模型(LLM)被用作许多领域自主决策与规划的策略。尽管LLM具备强大的推理能力,但在长周期任务中表现不佳,尤其是在部分可观测条件下。世界模型是提升策略性能的有效途径,无论是在训练阶段还是推理阶段。在推理阶段,当前智能体使用世界模型模拟候选动作的后果再执行,这能优化决策。但我们认为,仅模拟不足以应对部分可观测下的决策:模拟无法充分捕捉当前状态的不确定性,而这对智能体的准确决策至关重要。我们提出基于信念的世界模型(BB-WMs),该模型可建模并维护一种信念,LLM可查询此信念以获取当前状态的已知信息与不确定信息。在开发学习准确BB-WMs的方法前,我们先提出一个更基础的问题:将世界模型的信念直接暴露给LLM策略是否能提升决策效果?我们的结果表明,让LLM智能体获取世界模型的信念,可提升部分可观测条件下的任务性能,且与现有基于模拟的世界模型互补。代码已发布于该https URL。

英文摘要

Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning capabilities, LLMs struggle with long-horizon tasks, especially under partial observability. World models are a promising way to enhance policy performance, both during training and inference. During inference, agents currently use world models to simulate the consequences of candidate actions before choosing an action, which can improve decision-making. However, we argue that simulation alone is an incomplete interface for decision-making under partial observability: simulation does not adequately capture uncertainty about the current state, which agents may need for accurate decision-making. We address this limitation with Belief-Based World Models (BB-WMs), which maintain a belief that LLMs can query to access information on what is known and uncertain about the current state. Before developing methods to learn accurate BB-WMs, this paper focuses on a more fundamental question: does exposing a world model's belief directly to an LLM policy improve decision-making? Our results show that giving LLM agents access to beliefs improves task performance under partial observability, while remaining complementary to existing simulation-based world models. Code: https://github.com/skumar-ml/belief-world-models.

补充信息

↑