arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28416cs.CLcs.AIcs.LG

Agent-Editing World Model:重新思考LLM智能体的世界建模

Agent-Editing World Model: Rethinking World Modeling for LLM Agents

Shuang Sun, Guoxin Chen, Fanzhe Meng, Jia Deng, Huatong Song, Jinhao Jiang, Wayne Xin Zhao, Hongteng Xu, Ji-Rong Wen

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体任务状态污染问题,提出AEWM世界模型,通过Action Judge区分决策类型和State Revision编辑噪声延续,结合真实执行的EditAct提升性能,在多个基准上超越最强基线。

中文摘要 AI 辅助

大型语言模型(LLM)的最新进展使智能体能够在多样化的环境中处理长时程任务。为了进一步提升智能体性能,现有的语言世界模型通常预测环境观测,然而,在可获得真实反馈的情况下,重建高熵且依赖执行过程的工具响应价值有限。与此同时,智能体遭受“任务状态污染”问题,即未经证实的假设和过时的计划在历史中持续存在,从而扭曲后续决策。我们提出Agent-Editing World Model(AEWM),该模型建模推理和行动如何塑造未来任务进展,而非模拟工具响应。AEWM结合Action Judge以区分关键(Critical)、探索性(Exploratory)和噪声(Noisy)决策,并利用State Revision从相同的观测历史中编辑噪声的推理-行动延续。EditAct将这些能力与真实执行集成,直接改变后续决策所依据的状态,而非仅提供批评。我们通过中期训练和监督微调,在Search、Terminal和Software Engineering领域训练AEWM。AEWM在我们的Action Judge基准上达到70.5%的宏F1分数,超过最强前沿基线10.6个百分点。在六个基准和三个智能体骨干网络上,EditAct相比最强基线将平均分数提升3.2至6.7个百分点。此外,对经过验证的EditAct轨迹进行拒绝采样微调(称为AEWM-RFT),在没有在线AEWM指导的情况下,在三个领域上比Self-RFT提升2.2至2.6个百分点。

英文摘要

Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available. Meanwhile, agents suffer from \emph{task-state contamination}, where unsupported assumptions and outdated plans persist in history and distort subsequent decisions. We propose the \textbf{Agent-Editing World Model (AEWM)}, which models how reasoning and actions shape future task progress rather than simulating tool responses. AEWM combines \textbf{Action Judge} to distinguish \textsc{Critical}, \textsc{Exploratory}, and \textsc{Noisy} decisions with \textbf{State Revision} to edit noisy reasoning--action continuations from the same observed history. \textbf{EditAct} integrates these capabilities with real execution, directly changing the state underlying subsequent decisions rather than merely providing critiques. We train AEWM across Search, Terminal, and Software Engineering through mid-training and supervised fine-tuning. AEWM achieves 70.5\% macro-F1 on our Action Judge benchmark, exceeding the strongest frontier baseline by 10.6 points. Across six benchmarks and three agent backbones, EditAct improves average scores by 3.2--6.7 points over the strongest baseline. Furthermore, rejection sampling fine-tuning on verified EditAct trajectories, termed \textbf{AEWM-RFT}, improves over Self-RFT by 2.2--2.6 points across three domains without online AEWM guidance.

发表机构

  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

↑