arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从轨迹到智能体世界:用于交互式环境模拟的智能体语言世界模型

From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation

Quanyu Long, Xiao Chen, Jianda Chen, Haozhen Zhang, Qisheng Hu, Jianzhu Bao, Wenya Wang

arXiv 2610.06100首次发表:更新:

发表机构

Nanyang Technological University; The Hong Kong Polytechnic University(南洋理工大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出智能体语言世界建模范式及Trace2Env框架,利用历史交互轨迹构建环境世界手册,由世界模型智能体模拟环境,提升LLM智能体训练与评估中环境模拟的保真度和一致性。

AI 中文摘要

真实环境副本对于训练和评估LLM智能体越来越有价值,然而原始系统可能无法访问或难以复现。我们探索了智能体语言世界建模:与其重建一个可执行环境,不如让一个世界模型智能体充当任务智能体的环境,并支持忠实且有状态的模拟。我们通过Trace2Env实例化了这一范式,这是一个无需学习的框架,适用于原始系统不可用但历史交互轨迹仍可访问的场景。Trace2Env将这些轨迹重建为可复用的环境世界手册,其中包含环境模式、基于证据的事实和归纳出的行为知识。在运行时,世界模型智能体主动查阅世界手册以及持续的 episodic 状态,以推断每个动作的观察结果和持久的状态影响。在九个环境中,Trace2Env在下一观测保真度和长时程交互一致性方面均优于传统的基于提示的LWM。在多轮交互中,针对Trace2Env生成的任务智能体动作在真实环境中重放时更常保持有效,表明其模拟动态在连续轮次中更好地保留了早期动作的后果。这些结果确立了智能体语言世界建模作为在不重建原始可执行系统的情况下构建真实环境副本的替代方向。

英文摘要

Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language world modeling: rather than rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation. We instantiate this paradigm with Trace2Env, a learning-free framework for settings where the original system is unavailable but historical interaction traces remain accessible. Trace2Env reconstructs these traces into a reusable environment worldbook containing environment schemas, grounded evidence, and induced behavioral knowledge. At runtime, the world model agent actively consults the worldbook together with persistent episodic state to infer each action's observation and lasting state effects. Across nine environments, Trace2Env improves both next-observation fidelity and long-horizon interaction consistency over conventional prompt-based LWMs. In multi-turn interaction, task agent actions generated against Trace2Env remain valid more often when replayed in the real environment, indicating that its simulated dynamics better preserve the consequences of earlier actions across successive turns. These results establish agentic language world modeling as an alternative direction for building realistic environment replicas without reconstructing the original executable system.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑