arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06637cs.CL

通过结构化推理进行长视野文本世界建模

Long-Horizon Textual World Modeling through Structured Reasoning

  • University of Illinois Chicago(伊利诺伊大学芝加哥分校)
  • Intuit AI Research(Intuit AI 研究院)

机构由 AI 辅助整理,请以论文原文为准。

Fangxin Wang, Xiang Gao, Yuguang Yao, Kaiwen Dong, Nikash Walia, Kamalika Das

AI总结:

提出将多步转移建模为文本状态的结构化推理,以稀疏变化、预测增益和中间奖励解决长视野预测误差累积,在多个基准上取得最优性能。

AI中文摘要:

世界模型必须预测环境在动作序列下如何演变,使智能体能够在行动之前比较可能的未来并对反事实动作进行推理。长视野预测通常通过递归应用单步转移模型获得,但中间误差会随时间累积。多步动力学模型则直接以未来动作序列为条件并预测其后果,但随着视野增长,学习难度加大:模型必须跟踪轨迹中相互作用的状态变化,端点监督提供较弱的信用分配,且中间预测可能保持合理但丢失后续状态所需的信息。我们表明,这些挑战可以通过将多步转移的内部演变转化为对文本世界状态的结构化推理来解决:对稀疏状态变化的推理减轻了状态跟踪的负担,预测增益目标奖励学习到的状态优于匹配的以原始历史为条件的预测器,中间预测奖励沿轨迹监督每个状态。由于这些中间状态是世界的显式文本表示,它们提供了语义上有意义的目标,可在训练期间进行检查、评分和纠正。在ScienceWorld、Jericho和CEO-Bench上,我们的方法相对于直接以原始历史为条件的递归和非递归基线,实现了最强的平均长视野性能,且增益在更长视野下增加。在受控的反事实研究中,我们的模型也是唯一对未来动作具有统计显著敏感性的模型。

英文摘要:

World models must predict how an environment evolves under sequences of actions, enabling agents to compare possible futures and reason about counterfactual actions before acting. Long-horizon prediction is commonly obtained by recursively applying a one-step transition model, but intermediate errors can compound over time. Multi-step dynamics models instead condition on a sequence of future actions and predict their consequences directly, but become harder to learn as horizon grows: the model must track interacting state changes across the trajectory, endpoint supervision provides weak credit assignment, and intermediate predictions can remain plausible while losing information needed for later states. We show that these challenges can be addressed by casting the internal evolution of a multi-step transition as structured reasoning over textual world states: reasoning over sparse state changes reduces the burden of state tracking, a predictive-gain objective rewards the learned state for improving over a matched predictor that conditions on raw history instead, and intermediate predictive rewards supervise each state along the trajectory. Because these intermediate states are explicit textual representations of the world, they provide semantically meaningful targets that can be inspected, scored, and corrected during training. Across ScienceWorld, Jericho, and CEO-Bench, our approach achieves the strongest average long-horizon performance against recursive and non-recursive baselines that condition directly on raw history, with gains increasing at longer horizons. In a controlled counterfactual study, our model is also the only one with statistically significant sensitivity to future actions.

↑