arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-05-18 至 2026-05-18 共收录 14 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 14 篇

2605.15618 2026-05-18 cs.CV cs.AI 94%

Latent Video Prediction Learns Better World Models

潜在视频预测学习更好的世界模型

Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam, James Bailey, Naveed Akhtar

机构 * The University of Melbourne(墨尔本大学) Monash University(莫纳什大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文系统研究了潜在预测模型在世界模型中的鲁棒性,发现其在特征可区分性、抗污损性、细粒度辨别、遮挡鲁棒性和时间方向敏感性等方面表现优异,优于其他视频基础模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15725 2026-05-18 cs.CV cs.AI cs.RO 94%

DiLA: Disentangled Latent Action World Models

DiLA:解耦的潜在动作世界模型

Tianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang, Si Wu

机构 * Peking-Tsinghua Center for Life Sciences, Academy for Advanced Interdisciplinary Studies, IDG/McGovern Institute for Brain Research, Peking University(北京大学-清华生命科学中心,先进跨学科研究院,IDG/麦克戈文脑科学研究院,北京大学) Center of Quantitative Biology, Peking University(北京大学定量生物学中心) School of Psychological and Cognitive Sciences, Key Laboratory of Machine Perception (Ministry of Education), Peking University(心理与认知科学学院,机器感知重点实验室(教育部),北京大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 DiLA通过内容-结构解耦解决动作抽象与生成保真度的平衡问题,实现高质量视频生成和动作迁移。

Comments Project Page: http://disentangled-latent-action-world-models.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15967 2026-05-18 cs.AI cs.CV cs.LO 94%

Deterministic Event-Graph Substrates as World Models for Counterfactual Reasoning

确定性事件-图子结构作为世界模型用于反事实推理

Fabio Rovai

机构 * Tesseract Academy(Tesseract学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出事件图子结构作为世界模型,通过结构化干预词汇fork日志来回答反事实查询,证明了解释性与反事实性查询的对偶性,并在CLEVRER验证规模上评估了基于领域无关子结构运行时的C++解释器。

Comments 10 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15391 2026-05-18 cs.CV cs.AI 94%

PanoWorld: Geometry-Consistent Panoramic Video World Modeling

PanoWorld:几何一致的全景视频世界建模

Le Jiang, Xiangyu Bai, Bishoy Galoaa, Shayda Moezzi, Caleb James Lee, Tooba Imtiaz, Edmund Yeh, Jennifer Dy, Yanzhi Wang, Sarah Ostadabbas

机构 * Northeastern University(东北大学)

专题命中 通用世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 PanoWorld通过几何和动态一致性建模生成一致的360度视频,提升了空间理解能力,适用于具身AI应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15256 2026-05-18 cs.CV 93%

ReactiveGWM: Steering NPC in Reactive Game World Models

ReactiveGWM:引导NPC在反应式游戏世界模型中

Zeqing Wang, Danze Chen, Zhaohu Xing, Zizhao Tong, Yinhan Zhang, Xingyi Yang, Yeying Jin

机构 * Tencent(腾讯) National University of Singapore(新加坡国立大学) The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 ReactiveGWM通过解耦玩家控制与NPC行为,实现动态交互合成,使NPC策略能跨游戏迁移,无需重新训练即可实现可控的NPC互动。

Comments The code is available at https://inv-wzq.github.io/ReactiveGWM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15465 2026-05-18 cs.LG eess.SP 92%

Toward World Modeling of Physiological Signals with Chaos-Theoretic Balancing and Latent Dynamics

向生理信号的世界建模迈进:基于混沌理论的平衡与潜在动态

Yunfei Luo, Xi Chen, Yuliang Chen, Lanshuang Zhang, Md Mofijul Islam, Siwei Zhao, Peter Kotanko, Subhasis Dasgupta, Andrew Campbell, Rakesh Malhotra, Tauhidur Rahman

机构 * University of California San Diego(加州大学圣地亚哥分校) Dartmouth College(达特茅斯学院) Amazon Web Services(亚马逊网络服务) Sanderling Renal Services(Sanderling 肾脏服务) Renal Research Institute(肾脏研究研究所) Icahn School of Medicine at Mount Sinai(辛克尔医学院(Mount Sinai))

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);latent dynamics(title);world models(abstract)

AI总结 本文提出NormWear-2模型,通过将多变量生理信号与临床干预变量编码到共享潜在空间,结合先验知识推理与非参数潜在状态转移适应,实现多时间尺度的预测。混沌理论平衡动态制度多样性提升了表示鲁棒性,且在不同临床场景下表现优异。

Comments NormWear Collection: https://huggingface.co/collections/mosaic-laboratory/normwear

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15733 2026-05-18 cs.NE cs.AI cs.CV 90%

Structure Abstraction and Generalization in a Hippocampal-Entorhinal Inspired World Model

在启发式世界模型中的结构抽象与泛化

Tianqiu Zhang, Muyang Lyu, Xiao Liu, Si Wu

机构 * Peking-Tsinghua Center for Life Sciences, Academy for Advanced Interdisciplinary Studies, IDG/McGovern Institute for Brain Research, Center of Quantitative Biology, School of Psychological and Cognitive Sciences, Key Laboratory of Machine Perception (Ministry of Education), Peking University(北京大学-清华大学生命科学中心,先进跨学科研究院,IDG/麦克戈文脑科学研究院,定量生物学中心,心理与认知科学学院,机器感知重点实验室(教育部),北京大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出了一种脑启发的分层模型,通过逆向模型提取潜在转换并构建预测视觉世界模型,展示了在连续高维动态中同时提取抽象结构的能力,实现了结构泛化。

Comments Project page: https://hpc-mec-worldmodel.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15705 2026-05-18 cs.RO cs.AI 90%

Feedback World Model Enables Precise Guidance of Diffusion Policy

反馈世界模型使扩散策略获得精准指导

Tuo An, Jindou Jia, Gen Li, Jingliang Li, Chuhao Zhou, Pengfei Liu, Bofan Lyu, Jiaqi Bai, Xinying Guo, Geng Li, Jianfei Yang

机构 * MARS Lab, Nanyang Technological University(南洋理工大学MARS实验室)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出反馈世界模型,通过实时反馈修正预测误差,提升机器人决策性能,实验显示在分布偏移下预测准确率和策略表现显著提升。

Comments 21 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08398 2026-05-18 cs.CV 90%

VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?

VideoVerse: 你的T2V生成器有世界模型能力来合成视频吗?

Zeqing Wang, Xinyu Wei, Bairui Li, Zhen Guo, Jinrui Zhang, Hongyang Wei, Keze Wang, Lei Zhang

机构 * Sun Yat-sen University(中山大学) Hong Kong Polytechnic University(香港理工大学) Tsinghua University(清华大学) OPPO Research Institute(OPPO研究院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 VideoVerse通过评估T2V模型对复杂时间因果关系和世界知识的理解能力,揭示现有模型与理想世界建模能力的差距。

Comments 26 Pages, 10 Figures, 14 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16154 2026-05-18 cs.LG cs.RO 82%

Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking

学习结果分歧之处:通过概率块掩码实现高效的VLA强化学习

Vaidehi Bagaria, Nikshep Grampurohit, Pulkit Verma

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出概率块掩码(PCM),通过选择性分配梯度计算来提升GRPO-based VLA RL的效率,实现更快的训练速度和更低的内存消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15843 2026-05-18 cs.CV 81%

WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes

WorldAct:将单体3D世界激活为可交互的以对象为中心的场景

Jichen Hu, Jiawei Guo, Jiazhong Cen, Chen Yang, Sikuang Li, Wei Shen

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Inc(华为公司)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 WorldAct通过多模态代理将静态生成的3D世界分解为可编辑的交互场景,支持对象级编辑和任务执行,保留全局一致性。

Comments Project page: https://sjtu-deepvisionlab.github.io/WorldAct

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01970 2026-05-18 cs.AI cs.LG 69%

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models

小规模可泛化提示预测模型可引导大推理模型的高效强化学习后训练

Yun Qu, Qi Wang, Yixiu Mao, Heming Zou, Yuhang Jiang, Weijie Liu, Clive Bai, Kai Yang, Yangkun Chen, Saiyong Yang, Xiangyang Ji

机构 * Department of Automation, Tsinghua University, Beijing, China(自动化系,清华大学,北京,中国) LLM Department, Tencent, Beijing, China(大模型部门,腾讯,北京,中国)

专题命中 通用世界模型 :predictive model(title,abstract);predictive models(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出GPS方法,通过轻量级生成模型进行提示难度的贝叶斯推断,结合中间难度优先和历史锚定多样性,提升大模型强化学习后的训练效率和测试效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16205 2026-05-18 cs.AI cs.CL cs.LG cs.MA cs.SY eess.SY 60%

Context, Reasoning, and Hierarchy: A Cost-Performance Study of Compound LLM Agent Design in an Adversarial POMDP

上下文、推理与层次:在对抗性POMDP中的复合LLM代理设计成本-性能研究

Igor Bogdanov, Chung-Horng Lung, Thomas Kunz, Jie Gao, Adrian Taylor, Marzia Zaman

机构 * Carleton University(卡尔顿大学)

专题命中 通用世界模型 :environment model(abstract);分类 cs.AI、cs.LG、cs.MA

AI总结 研究探讨了在对抗性部分可观测序贯环境中,复合LLM代理设计的上下文、推理和层次分解对性能与成本的影响,发现程序化状态抽象在成本效率上表现最佳,而分层分解无需推理可获得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15461 2026-05-18 cs.LG cs.AI 50%

DrugSAGE:Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery

DrugSAGE: 自演化代理经验用于高效前沿药物发现

Yikun Zhang, Xiwei Cheng, Tianyu Liu, Yuanqi Du, Wengong Jin

机构 * Northeastern University(东北大学) Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所) Yale University(耶鲁大学) Microsoft Research New England(微软研究院新英格兰分部)

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 DrugSAGE通过自演化代理经验框架,高效构建前沿药物发现模型,跨任务记忆提升模型性能,实现零次搜索下的显著优势。

详情

展开后加载摘要…

URL PDF HTML 收藏