arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OPINE-World:基于本体错误优先的交互式探索的程序化世界建模

OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration for ARC-AGI-3

David Courtis, Wenhao Li, Scott Sanner

arXiv 2607.01531首次发表:更新:

发表机构

University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出OPINE-World,一种在线交互学习面向对象的程序化世界模型的LLM智能体,通过本体错误度量引导探索,在ARC-AGI-3基准上无需逐游戏训练即解决20/25个游戏,动作效率达78.4。

AI 中文摘要

从交互中学习环境的行为是构建适应不熟悉任务的智能体的核心。基于深度网络的世界模型灵活但数据需求大且迁移性差。由LLM编写为源代码并通过反例引导归纳合成(CEGIS)精炼的程序合成世界模型数据高效且可重用,但主要应用于具有给定对象词汇的结构化状态世界,且单一程序搜索无法扩展到需要灵活假设对象结构的像素渲染环境。我们提出OPINE-World,一个从交互中在线学习面向对象的程序化世界模型的LLM智能体。OPINE-World在假设与测试循环中耦合两个协作智能体,一个在环境中行动,一个以代码形式合成模型并进行重放验证和基于模型的规划,并通过我们称为本体错误的贝叶斯对象类型充分性度量引导探索。我们在ARC-AGI-3(一个技能获取效率基准,其中对象词汇、目标和动作语义均被隐藏)上评估OPINE-World。OPINE-World无需逐游戏训练即解决了25个游戏中的20个,并达到78.4的动作效率分数(以人类基线为基准)。

英文摘要

Learning a useful world model from minimal interaction is central to building agents that adapt to unfamiliar tasks. Programmatic world modeling approaches such as WorldCoder quickly learn transition models when given pre-supplied symbolic representations; however, such approaches are insufficient when the symbolic representation and goal are unknown and continually changing, and are intractable under large action spaces due to rigid planning heuristics. Extending the programmatic world modeling approach, OPINE-World additionally maintains an abstracted representation in which provisional objects, incomplete causal explanations, and unresolved interactions can exist in a Bayesian, exploration-centric hypothesis space. This enables learning an unknown, dynamically evolving state, transition, and goal. Exploration and program revision share this evolving abstraction, allowing new evidence to revise both the hypothesized mechanics and their implementation in tandem at test time. ARC-AGI-3 presents this open-world problem as a general-intelligence learning benchmark. Through our OPINE-world model learning paradigm, we improve the Relative Human Action Efficiency score of Claude Opus 4.8 high from 1.5% to 78.40% on the ARC-AGI-3 benchmark.

CommentsUnder review at ICLR 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑