arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HyperWorld:超图结构化状态序列化改进学习型文本世界模型

HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models

Yun-Jian Zhang, Chen-Wei Liang, Tian-Yi Zhang, Jian Ding, Yi-Lun Wu, Ao-Bo Li, Wei-Cong Su, Saifullah, Hong-Yu An, Mu-Jiang-Shan Wang

arXiv 2609.00002首次发表:更新:

发表机构

Shenzhen Kaihong Digital Industry Development Co., Ltd.; University of New South Wales; Zhejiang University; University of Macau; Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳市凯虹数字工业发展有限公司; 新南威尔士大学; 浙江大学; 澳门大学; 中国科学院深圳先进技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HyperWorld研究显示,超图结构化状态序列化作为归纳偏置,可提升学习型文本世界模型在分布偏移、中小规模模型下的规划与预测性能,优于独立句子、成对三元组等序列化方式。

AI 中文摘要

世界模型使语言模型智能体能够预测环境动态并在行动前进行规划。在文本环境中,模型必须从序列化状态描述中学习符号化的行动效果,但序列化结构的作用尚未得到充分探索。我们提出HyperWorld,这是一项针对学习型文本世界模型的状态序列化的对照研究。我们将原始观测结果与同一真实状态的三种符号化序列化方式进行对比:独立句子、成对三元组,以及围绕实体和关系对多个相关事实进行分组的以实体为中心的超边单元。所有变体均采用相同的训练目标:给定一个状态和一个行动,预测符号化效果或判定该行动不可行。在不同模型规模、数据预算、分布内和分布外测试环境中,超边序列化在0.5B至1.5B模型及分布偏移场景下带来的提升最为显著。更大的模型缩小了差距,成对三元组在分布内精确匹配上可与超边相当或略微超过超边,但超边在分布外事实F1值上表现最佳,且在可行性检测与效果预测之间实现了最佳的中小规模权衡。在下游贪心规划中,超边世界模型在 tested 表示中也达到了最高成功率。这些结果表明,高阶状态组织是学习型符号世界模型的一种简单但有效的归纳偏置,尤其在模型容量有限或测试环境与训练环境存在差异时效果明显。

英文摘要

World models enable language-model agents to predict environment dynamics and plan before acting. In text environments, the model must learn symbolic action effects from serialized state descriptions, but the role of serialization structure remains underexplored. We present HyperWorld, a controlled study of state serialization for learned textual world models. We compare raw observations with three symbolic serializations of the same ground-truth state: independent sentences, pairwise triples, and entity-centered hyperedge units that group multiple related facts around entities and relations. All variants use the same training objective: given a state and an action, predict symbolic effects or judge the action infeasible. Across model scales, data budgets, and in-distribution and out-of-distribution test worlds, hyperedge serialization gives the clearest gains for 0.5B--1.5B models and under distribution shift. Larger models reduce the gap, and pairwise triples can match or slightly exceed hyperedges on in-distribution exact match, but hyperedges achieve the strongest out-of-distribution fact F1 and the best small-to-medium scale trade-off between feasibility detection and effect prediction. In downstream greedy planning, the hyperedge world model also attains the highest success rate among the tested representations. These results show that higher-order state organization is a simple but effective inductive bias for learned symbolic world models, especially when model capacity is limited or test environments differ from training.

Comments10 pages, 4 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑