arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-02-12 至 2026-02-12 共收录 9 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 9 篇

2602.08971 2026-02-12 cs.CV cs.RO 96%

WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models

WorldArena: 一个用于评估具身世界模型感知与功能效用的统一基准

Yu Shang, Zhuohang Li, Yiding Ma, Weikang Su, Xin Jin, Ziyou Wang, Lei Jin, Xin Zhang, Yinzhou Tang, Haisheng Su, Chen Gao, Wei Wu, Xihui Liu, Dhruv Shah, Zhaoxiang Zhang, Zhibo Chen, Jun Zhu, Yonghong Tian, Tat-Seng Chua, Wenwu Zhu, Yong Li

机构 * Tsinghua University, Beijing, China(清华大学) Shanghai Jiao Tong University, Shanghai, China(上海交通大学) The University of Hong Kong, Hong Kong SAR, China(香港大学) Princeton University, Princeton, NJ, USA(普林斯顿大学) Chinese Academy of Sciences, Beijing, China(中国科学院) University of Science(科学技术大学) Peking University, Beijing, China(北京大学) National University of Singapore, Singapore(新加坡国立大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title,abstract);world model(title,abstract)

AI总结 WorldArena是一个统一的基准,用于评估具身世界模型的感知和功能效用,揭示高视觉质量与强具身任务能力之间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10717 2026-02-12 cs.RO 95%

Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation

说、梦、做:学习视频世界模型以驱动指令式机器人操作

Songen Gu, Yunuo Cai, Tianyu Wang, Simo Wu, Yanwei Fu

机构 * Fudan University(复旦大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title);world model(title,abstract)

AI总结 本文提出一种视频条件动作框架,通过生成稳健的视频模型和对抗性蒸馏,提升机器人操作中的预测能力和空间准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08025 2026-02-12 cs.CV cs.AI 94%

MIND: Benchmarking Memory Consistency and Action Control in World Models

MIND:世界模型中内存一致性与动作控制的基准测试

Yixuan Ye, Xuanyu Lu, Yuxin Jiang, Yuchao Gu, Rui Zhao, Qiwei Liang, Jiachun Pan, Fengda Zhang, Weijia Wu, Alex Jinpeng Wang

机构 * CSU-JPG, Central South University(中南大学) National University of Singapore(新加坡国立大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Nanyang Technological University(南洋理工大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 MIND提出首个开放域闭环基准,评估世界模型的内存一致性和动作控制能力,揭示当前模型在长期记忆保持和动作空间泛化上的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22130 2026-02-12 cs.AI cs.SE 93%

World of Workflows: A Benchmark for Bringing World Models to Enterprise Systems

世界工作流:一个将世界模型引入企业系统的基准

Lakshya Gupta, Litao Li, Yizhe Liu, Sriram Ganapathi Subramanian, Kaheer Suleman, Zichen Zhang, Haoye Lu, Sumit Pasupalak

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 World of Workflows提出一个基于ServiceNow的企业系统基准,揭示前沿LLMs在动态建模中的盲区,并强调基于现实世界建模的可靠性需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11021 2026-02-12 cs.RO cs.AI cs.CV 91%

ContactGaussian-WM: Learning Physics-Grounded World Model from Videos

ContactGaussian-WM: 从视频中学习物理基础的世界模型

Meizhong Wang, Wanxin Jin, Kun Cao, Lihua Xie, Yiguang Hong

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 ContactGaussian-WM通过从稀疏视频中学习物理定律,提升机器人规划与模拟中的复杂动态场景建模能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10390 2026-02-12 cs.LG cs.AI 90%

Affordances Enable Partial World Modeling with LLMs

affordances 使 LLMs 能实现部分世界建模

Khimya Khetarpal, Gheorghe Comanici, Jonathan Richens, Jeremy Shar, Fei Xia, Laurent Orseau, Aleksandra Faust, Doina Precup

机构 * Google Deepmind(谷歌DeepMind)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出利用 affordances 构建部分世界模型,通过在多任务中引入分布稳健的 affordances,提高搜索效率并提升奖励表现。

Comments 18 pages, 5 figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10884 2026-02-12 cs.CV 90%

ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving

ResWorld: 用于端到端自动驾驶的时序残差世界模型

Jinqing Zhang, Zehua Fu, Zelin Xu, Wenying Dai, Qingjie Liu, Yunhong Wang

机构 * State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,北京航空航天大学) Zhongguancun Laboratory, Beijing, China(中关村实验室) Beijing Jingwei Hirain Technologies Co., Inc.(北京京wei Hirain科技有限公司) Hangzhou Innovation Institute, Beihang University, Hangzhou, China(杭州创新研究院,北京航空航天大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 ResWorld通过时序残差世界模型和未来引导轨迹细化模块,提升端到端自动驾驶的规划性能。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04345 2026-02-12 cs.SD cs.AI cs.LG 71%

AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds

AUDETER:一种大规模深度伪造音频检测数据集用于开放世界

Qizhou Wang, Hanxun Huang, Guansong Pang, Sarah Erfani, Christopher Leckie

机构 * The University of Melbourne(墨尔本大学) Singapore Management University(新加坡管理学院)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG

AI总结 AUDETER是一种大规模深度伪造音频检测数据集,通过课程学习方法提升开放世界检测性能,实现1.87%的EER在野外测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07593 2026-02-12 stat.ML cs.LG stat.ME 69%

Diffusion posterior sampling for simulation-based inference in tall data settings

扩散后验采样用于高数据设置下的基于模拟的推断

Julia Linhart, Gabriel Victorino Cardoso, Alexandre Gramfort, Sylvain Le Corff, Pedro L. C. Rodrigues

机构 * Université Paris-Saclay(巴黎-萨克雷大学) Inria(法国国家信息与自动化技术研究院) CEA(法国原子能委员会) CMAP, École Polytechnique(高等理工学院CMAP部门) Institut Polytechnique de Paris(巴黎理工 institute) CNRS(法国国家科学研究中心) Grenoble INP(格勒诺布尔INP) LJK(格勒诺布尔联合实验室)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 本文提出了一种无需兰格-动态步骤的扩散后验采样方法,提高了高数据设置下基于模拟的推断效率和稳定性。

Comments 49 pages, 24 figures, 3 tables, 2 algorithms, 12 appendices, TMLR acceptance

详情

展开后加载摘要…

URL PDF HTML 收藏