arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01861cs.AI

信念校准优化:智能体优化的显式世界模型

Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization

Yuhan Chen, Zhihua Tian, Mahavir Dabas, Charith Peris, Rahul Gupta, Ming Jin, Feiyang Kang, Siyuan Zhang, Nan Wang, Ruoxi Jia

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出信念校准优化(BCO)方法,将智能体的隐式信念转化为显式持久上下文文档作为世界模型,在5类基准测试中提升了LLM智能体的训练通过率,且其内容具备可复用价值。

中文摘要 AI 辅助

大语言模型(LLM)智能体的性能取决于冻结模型周围的支撑框架。提升该框架的常用方式是将编码智能体用作优化器:它读取当前得分与轨迹,迭代编辑源代码,每轮生成一个新候选方案。每次编辑的选择依据是对环境响应的信念:哪里出错了,哪项修改会有帮助。这种信念通常是隐式的,存在于编码智能体对当前调用的推理中,或潜藏在其参数中,而非以显性形式存在。后续调用虽能获取得分与轨迹,却无法利用该信念。本文提出信念校准优化(Belief-Calibrated Optimization, BCO)方法,将上述信念以持久的上下文文档形式记录下来,并在评估新候选方案时持续修订该文档。生成的文档即世界模型,是当前环境对编辑的响应情况的记录。将BCO加入标准循环后,在涵盖记忆问答、工具使用问答、代码即行动应用智能体及终端智能体的5个基准测试中,BCO达到的训练通过率高于仅缺少世界模型的匹配对照组。在未用于候选方案选择的所有保留拆分测试中,该差距依然存在。在目标模型替换(冻结模型被替换但支撑框架不变)后,所选BCO支撑框架在测试任务中表现领先,仅在上下文窗口溢出导致未完成的情况下例外。随后的离线消融实验探究该差距是否源于世界模型的内容:与未提供文档或提供内容被篡改的同格式副本的预测器相比,基于累积文档的新预测器能更准确地预测环境响应。对比结果表明,该文档的内容承载了可复用信息,而非仅形式上的作用。

英文摘要

The performance of an LLM agent depends on the scaffold around a frozen model. A common way to improve that scaffold is to use a coding agent as an optimizer: it reads current scores and traces and iteratively edits the source, producing a new candidate each round. Each edit is chosen according to a belief about how the environment will respond: what went wrong, and which change should help. That belief is typically implicit. It lives in the coding agent's reasoning on the current call, or remains latent in its parameters, rather than as something written down. Later calls therefore see scores and traces, but they do not use that belief. We introduce Belief-Calibrated Optimization (BCO), a method that writes that belief down as a persistent in-context document and continually revises that document as new candidates are evaluated. The resulting document is a world model: the current account of how the environment responds to edits. Added to an otherwise standard loop, BCO reaches a higher train passrate than a matched control that lacks only the world model, on five benchmarks spanning memory QA, tool-use QA, code-as-action app agents, and terminal agents. The gap remains on every held-out split, which is not used to select the candidate. After a target-model swap, in which the frozen model is replaced and the scaffold is not, the selected BCO scaffold leads on the tasks we test, except where context-window overruns leave it unfinished. An offline ablation then asks whether that gap comes from what the world model says. A fresh predictor given the accumulated document forecasts how the environment will respond more accurately than predictors given either no document or a same-form copy whose content has been falsified. The comparison indicates that the document carries reusable information in its content, not only in its form.

发表机构

  • Virginia Tech(弗吉尼亚理工大学)
  • Amazon(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

↑