arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ALDER:通过在世界中行动来发现世界规律

ALDER: Discovering the Laws of a World by Acting in It

Teng Cao, Yu Deng, Quentin Delfosse, Kristian Kersting

arXiv 2609.33728首次发表:更新:

发表机构

Technische Universität Darmstadt; Intrinsic; German Research Center for AI; The Hessian Center for AI(达姆施塔特工业大学; Intrinsic公司; 德国人工智能研究中心; 黑森人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ALDER通过主动设计实验来测试和修订显式方程世界模型,能发现初始假设之外的规律,并用于目标导向控制。

AI 中文摘要

可靠的世界模型不仅应预测未来状态,还应以明确、透明且可测试的形式(如方程)表达行动如何改变世界。然而,依赖固定轨迹集的方法无法区分同样优秀的竞争性假设,而对固定预定义候选集进行搜索的方法则无法发现初始假设空间之外的方程。我们引入了ALDER(行动引导的规律发现、评估与修订),一种主动提出新颖实验以测试和修订模型的方法。具体而言,ALDER提出参数方程;数值优化器拟合其系数;独立验证器在留出数据上测试这些候选。为区分竞争性的有效假设,一个考虑成本和安全性的选择器以新颖实验的形式查询干预。由此产生的反例更新证据账本并指导下一次结构修订,而不兼容的规律则被丢弃。在内部基准、常微分方程方程发现任务和机器人实验中,ALDER发现了其初始公式集之外的规律,修复了失败的模型提议,以更少的交互区分了固定候选模型,并改善了分布外预测。此外,给定当前状态和目标,ALDER通过求解其验证过的世界模型定义的逆问题来选择控制行动。这些结果共同表明,基于显式方程的世界模型可以通过交互进行测试和修订,然后自然地用于指导目标导向的控制。

英文摘要

Reliable world models should not only predict future states but express how actions change the world in an explicit, transparent and testable form, such as equations. Yet methods that rely on a fixed set of trajectories cannot distinguish equally good competing hypotheses, while searches over a fixed set of predefined candidates cannot discover equations outside the initial hypothesis space. We introduce ALDER (Action-guided Law Discovery, Evaluation, and Revision), a method that actively proposes novel experiments to test and revise models. Specifically, ALDER proposes parametric equations; a numerical optimizer fits their coefficients; an independent verifier tests these candidates on held-out data. To distinguish between competing valid hypotheses, a cost- and safety-aware selector queries interventions, in the form of novel experiments. The resulting counterexamples update the evidence ledger and guide the next structural revision, while incompatible laws are discarded. Across an in-house benchmark, ODE equation discovery tasks, and robotic experiments, ALDER discovers laws beyond its initial formula set, repairs failed model proposals, distinguishes fixed candidate models with fewer interactions, and improves out-of-distribution prediction. Furthermore, given a current state and a target, ALDER selects control actions by solving the inverse problem defined by its validated world model. Together, these results show that explicit equation-based world models can be tested and revised through interaction, then naturally used to guide goal-directed control.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑