超越记忆:利用显式信念状态驾驭长时程智能体
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States
查看机构详情
- Nankai University(南开大学)
- Alibaba Group(阿里巴巴集团)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出推理时框架PoS,通过构建并持续维护显式信念状态作为决策上下文,检测并恢复“信念陷阱”,在四个基准上以三种LLM骨干取得最优性能,为长时程上下文管理提供新基础。
中文摘要 AI 辅助
大型语言模型(LLM)智能体现在能够承担日益复杂的任务,但它们将交互历史组织成记忆的方式并不能确保对当前世界形成连贯的理解。我们提出PoS,一个推理时框架,它构建并持续维护显式信念状态作为智能体的决策上下文。每个信念结合了对当前世界状态的估计与未解决的任务需求,明确指出了智能体仍需学习和完成的内容。为了保持这种信念的可靠性和可操作性,PoS验证其一致性并监控任务进度以检测“信念陷阱”(Belief Trapping),即智能体在未朝目标取得有意义进展的情况下继续行动。随后,恢复策略根据陷阱模式和未解决任务需求的类型进行定制。在涵盖执行和诊断的四个基准上的实验表明,PoS在所有三个LLM骨干网络上均取得了每个基准的最高整体性能。消融实验证明了一致性验证和恢复的重要性,而上下文扩展实验则显示了对上下文增长的韧性。这些结果共同支持信念构建和持续维护作为超越历史保留与压缩的长时程上下文管理的基础。
英文摘要
Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.