arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24744cs.AI

世界状态生成器

World State Generator

Sungheon Jeong, Sanggeon Yun, Ryozo Masukawa, Haleh Alimohamadi, Mahdi Imani, Mohsen Imani

首次发表
浏览论文内容

中文总结 AI 辅助

提出世界状态生成器(WSG),将计划写为世界可检查状态,利用约22.6万条失败与修复轨迹训练,使计划在运行中顺应世界规则,在7个基准上提升近300亿参数开源模型成功率至专有模型水平。

中文摘要 AI 辅助

语言智能体通过计划和行动解决复杂任务。世界拒绝的单个步骤会使目标遥不可及,而智能体接下来的行动决定了任务的成败。提示式规划者恰恰在这一环节失败:他们用新措辞重写被拒绝的步骤,却遭遇同样的拒绝,并在未取得进展的情况下耗尽尝试预算。他们失败的原因在于计划从未与世界绑定,因此拒绝在计划中没有可依附的对象。世界是任务运行的环境,它有自己的规则、可允许的行动和约束。我们在7个领域构建合成世界,并从中提取训练数据。一个程序强制执行每个世界的规则并评估其目标,且每个世界只有在目标从其初始状态可达时才会被接纳。智能体在内部运行,并留下经过验证的失败记录,以及将运行带到世界认证状态的修复,总计约22.6万条轨迹的记录。基于此记录,我们训练了世界状态生成器(WSG),该模型将计划写为世界的可检查状态,并保持计划与其运行的世界对齐。这种对齐正是用语言编写的计划所缺乏的,因为其运行的世界具有物理限制、逻辑依赖和必需顺序,而语言从未说明这些,计划仅在状态失败时才遇到这些规则。WSG将失败视为世界所陈述的规则,并重写剩余状态以遵守该规则,从而使计划在运行过程中顺应世界。在7个公开基准上,WSG将两个接近300亿参数的开源模型的端到端成功率提升至超过提示式方法,并达到专有模型的水平。

英文摘要

Language agents solve complex tasks through plans and actions. A single step the world refuses puts the goal out of reach, and what the agent does next decides the task. Prompted planners fail at exactly this point, rewriting the refused step in new words, meeting the same refusal, and burning the attempt budget without moving. They fail because the plan was never tied to the world, so a refusal has nothing in the plan to attach to. A world is where a task runs, and it has its own rules, its own admissible actions, and its own constraints. We build synthetic worlds across 7 domains and extract training data from them. A program enforces each world's rules and grades its goal, and every world is admitted only if its goal is reachable from its initial state. Agents run inside and leave verified failures paired with repairs that carried the run to a state the world certified, a record of about 226K trajectories. On this record we train the World State Generator, a model that writes a plan as checkable states of the world and keeps that plan aligned with the world it runs in. That alignment is what a plan written in language lacks, since the world it runs in has physical limits, logical dependencies, and required orders the language never states, and the plan encounters these rules only when a state fails. WSG takes that failure as the rule the world has stated and rewrites the remaining states to obey it, so the plan bends to the world as the run goes on. Across 7 public benchmarks, WSG raises end-to-end success for two open models near 30B parameters over prompting and brings to the level of proprietary model.

发表机构

  • University of California, Irvine(加利福尼亚大学尔湾分校)
  • Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

↑