arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17419cs.AI

世界模型科学:长时程LLM智能体中的自组织临界性、弱混沌与亚稳态信念动力学

World Model Science: Self-Organized Criticality, Weak Chaos, and Metastable Belief Dynamics in Long-Horizon LLM Agents

Xinyuan Song, Zekun Cai

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过自组织临界性、弱混沌和亚稳态信念动力学三个视角,分析长时程LLM智能体的轨迹,提出基于轨迹级动力学诊断的世界模型科学方法,并在22项实验中验证其有效性。

中文摘要 AI 辅助

长时程LLM智能体必须在观察、动作、工具调用和中间信念的扩展序列中维持任务状态。我们通过三个动力学视角研究这些轨迹:自组织临界性、弱混沌和亚稳态信念动力学。我们的框架将智能体隐含状态与基准锚定状态对齐,并在显式零模型下测量应力积累、错误雪崩、时间依赖性、局部-全局失配、有界发散、信念盆地转换和有限尺寸标度。在涵盖受控谜题、工具使用、具身任务、多跳检索、通用助手推理和生命游戏的22项实验中,我们发现局部有效的动作可以在全局状态保真度失效后持续存在,应力可以触发突然崩溃,错误序列表现出长记忆,依赖深度改变传播机制,更大的时程支持更大的雪崩。同时,发散保持有界,信念状态表现出亚稳态而非完全混沌行为,关于普遍幂律、临界点或共享干预最优的更强主张未得到支持。这些结果表明,基于轨迹级动力学诊断而非仅终端奖励的智能体世界模型科学是可行的。

英文摘要

Long-horizon LLM agents must maintain task state across extended sequences of observations, actions, tool calls, and intermediate beliefs. We study these trajectories through three dynamical views: self-organized criticality, weak chaos, and metastable belief dynamics. Our framework aligns agent-implied states with benchmark-grounded states and measures stress accumulation, error avalanches, temporal dependence, local--global mismatch, bounded divergence, belief-basin transitions, and finite-size scaling under explicit null models. Across 22 experiments spanning controlled puzzles, tool use, embodied tasks, multi-hop retrieval, general-assistant reasoning, and Game of Life, we find that locally valid actions can persist after global state fidelity fails, stress can trigger abrupt collapse, error sequences exhibit long memory, dependency depth changes the propagation regime, and larger horizons support larger avalanches. At the same time, divergence remains bounded, belief states show metastable rather than fully chaotic behavior, and stronger claims of universal power laws, critical points, or shared intervention optima are not supported. These results suggest a science of agent world models based on trajectory-level dynamical diagnostics rather than terminal reward alone.

发表机构

  • Emory University(埃默里大学)
  • The University of Tokyo(东京大学)
  • LocationMind

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑