arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10036cs.AIcs.LGcs.RO

信念状态引擎:为部分可观测条件下的原则性规划增强大语言模型

Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability

Arnab Chattopadhayay, Debdipta Halder

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型智能体在部分可观测环境下易过早承诺、策略漂移的问题,提出信念状态引擎(BSE),通过外部维护贝叶斯后验并仅暴露给模型,实现原则性规划,在Tiger和攻击图任务上提升回报与决策一致性。

中文摘要 AI 辅助

大语言模型智能体在广泛的任务中能够产生流畅的动作序列,然而一旦环境变为部分可观测,它们就会以特有的方式失败。模糊的反馈促使它们过早地做出承诺。单一的信息性观测可能将其不确定性坍缩到错误的假设上。随着历史增长,策略发生漂移。我们将这些症状追溯到一个共同的结构性原因。通常部署的大语言模型智能体是一个以历史为条件的策略,对隐藏状态没有显式的信念。我们提出了一种架构上的修复方案。信念状态引擎(BSE)是置于大语言模型外部的推理模块。它维护给定POMDP(部分可观测马尔可夫决策过程)模型潜在状态上的贝叶斯后验,并在每个决策步骤仅将该后验暴露给大语言模型。原始的动作-观测日志不被展示。我们提出了一个最小四公理规范,说明信念一致的内部状态必须满足什么条件,并证明与大语言模型配对的BSE是底层POMDP诱导的信念MDP上的一个健全马尔可夫策略。因此,只要大语言模型从不接触原始历史,它就继承了经典POMDP理论的贝尔曼最优性保证。我们在Tiger POMDP和红队攻击图任务上评估了该架构,与六个基线进行比较:反应式大语言模型、思维链、ReAct、自然语言信念追踪器、QMDP和POMCP。在两个领域中,BSE增强的智能体提高了任务回报、信念校准和决策一致性。十项针对性的消融研究隔离了每个架构选择的贡献,并确认该效果并非特定于任何单一模型。代码、环境规范、提示模板和种子日志随本文提供。

英文摘要

Large language model agents produce fluent action sequences across a wide range of tasks, yet they fail in characteristic ways once the environment becomes partially observable. Ambiguous feedback pushes them into premature commitments. A single informative observation can collapse their uncertainty onto the wrong hypothesis. Policies drift as the history grows. We trace these symptoms to a common structural cause. An LLM agent, as commonly deployed, is a history-conditioned policy with no explicit belief over hidden state. We propose an architectural fix. The Belief-State Engine (BSE) is an inference module placed outside the LLM. It maintains a Bayesian posterior over the latent states of a given POMDP (Partially Observable Markov Decision Process) model, and at each decision step it exposes only that posterior to the LLM. The raw action-observation log is not shown. We set out a minimal four-axiom specification of what a belief-consistent internal state must satisfy, and prove that the LLM paired with the BSE is a sound Markov policy on the belief MDP induced by the underlying POMDP. It therefore inherits the Bellman optimality guarantees of classical POMDP theory, provided the LLM is never exposed to the raw history. We evaluate the architecture on the Tiger POMDP and a red-team attack-graph task, against six baselines: a reactive LLM, Chain-of-Thought, ReAct, a natural-language belief tracker, QMDP, and POMCP. Across both domains, the BSE-augmented agent improves task return, belief calibration, and decision consistency. Ten targeted ablations isolate the contribution of each architectural choice confirms that the effect is not specific to any one model. Code, environment specifications, prompt templates, and seed logs accompany this paper.

补充信息

↑