arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双前沿:智能体何时能信任其世界模型?

Dual-Frontier: When Can an Agent Trust Its World Model?

Huatai Zhu, Qiang Chen, Ziqian Kou, Wenhao Li, Fei Wang, Yichao Cao, Xiu Su, Yi Chen

arXiv 2609.26293首次发表:更新:

发表机构

Central South University; The Hong Kong University of Science and Technology; Xiangjiang Laboratory; University of Sydney; University of Science and Technology of China(中南大学; 香港科技大学; 湘江实验室; 悉尼大学; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对世界模型引导决策失败归因不可识别问题,提出双前沿原则,通过认证优势界决定是否采纳模型决策,否则进行模型验证,保证回报非递减并提升决策质量。

AI 中文摘要

学习到的世界模型正成为通用智能体的关键组成部分:通过预测行动后果,它们支持规划与决策,同时减少对昂贵试错的依赖。这种依赖产生了一个根本性的模糊性:当世界模型引导的决策失败时,仅凭轨迹可能无法揭示是智能体的决策规则还是世界模型导致了损失。我们将这一失败归因问题形式化为回报损失的反事实分解,并证明其组成部分无法从被动交互中识别,即使对于有限时域规划器也是如此。这一障碍促使我们提出双前沿(Dual-Frontier)学习原则,该原则仅当预测优势超过决策相关世界模型误差的认证界时,才允许世界模型引导的决策;否则,证据被分配给世界模型验证。行动条件价值界和闭环扩展保证了被采纳决策的非递减回报。校准门控和同时置信序列支持自适应证据重用,并具有充分且必要的验证界。受控学习模型实验验证了预测的失败模式和认证行为,而跨骨干工具使用基准在现实智能体世界模型流程中实例化了相同的先验证后提升规则,持续提高了决策质量和可靠性。

英文摘要

Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error. This reliance creates a fundamental ambiguity: when a world-model-guided decision fails, the trajectory alone may not reveal whether the agent's decision rule or the world model caused the loss. We formalize this failure-attribution problem as a counterfactual decomposition of return loss and prove that its components are not identifiable from passive interaction, even for finite-horizon planners. This obstruction motivates Dual-Frontier, a learning principle that admits a world-model-guided decision only when its predicted advantage exceeds a certified bound on decision-relevant world-model error; otherwise, evidence is allocated to world-model verification. Action-conditioned value bounds and a closed-loop extension guarantee non-decreasing return for admitted decisions. Calibrated gates and simultaneous confidence sequences support adaptive evidence reuse, with sufficient and necessary verification bounds. Controlled learned-model experiments validate the predicted failure modes and certification behavior, while cross-backbone tool-use benchmarks instantiate the same verify-then-promote rule in realistic agent world-model pipelines, consistently improving decision quality and reliability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑