arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-03-10 至 2026-03-10 共收录 19 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 19 篇

2601.01528 2026-03-10 cs.CV cs.AI cs.RO 96%

DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

DrivingGen:自主驾驶中生成视频世界模型的综合性基准

Yang Zhou, Hao Shao, Letian Wang, Zhuofan Zong, Hongsheng Li, Steven L. Waslander

机构 * University of Toronto(多伦多大学) CUHK MMLab(香港中文大学MMLab)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title);world model(title,abstract)

AI总结 DrivingGen提出首个综合性基准,用于评估生成驾驶世界模型的视觉真实性、轨迹合理性、时间一致性和可控性,揭示通用与专用模型间的权衡。

Comments ICLR 2026 Poster; Project Website: https://drivinggen-bench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07264 2026-03-10 cs.RO cs.AI 95%

Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving

具备运动学意识的潜在世界模型用于数据高效的自动驾驶

Jiazhuo Li, Linjiang Cao, Qi Liu, Xi Xiong

机构 * Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University(道路与交通工程重点实验室,教育部,同济大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出一种具备运动学意识的潜在世界模型框架,通过整合车辆运动学信息提升自动驾驶的样本效率和驾驶性能。

Comments 6 pages, 5 figures. Under review at IEEE ITSC

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19195 2026-03-10 cs.CV cs.AI 94%

Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks

重新思考驾驶世界模型作为感知任务的合成数据生成器

Kai Zeng, Zhanqian Wu, Kaixin Xiong, Xiaobao Wei, Xiangyu Guo, Zhenxin Zhu, Kalok Ho, Lijun Zhou, Bohan Zeng, Ming Lu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Wentao Zhang

机构 * Peking University(北京大学) Xiaomi EV(小米电动车) Huazhong University of Science and Technology(华中科技大学) Beijing Key Laboratory of Data Intelligence and Security (Peking University)(北京数据智能与安全重点实验室(北京大学)) Zhongguancun Academy(中关村学院)

专题命中 通用世界模型 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 Dream4Drive通过生成高质量的合成数据提升自动驾驶感知任务性能

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00296 2026-03-10 cs.RO cs.AI cs.CV cs.LG 94%

From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models

从像素到谓词:通过预训练视觉-语言模型学习符号世界模型

Ashay Athalye, Nishanth Kumar, Tom Silver, Yichao Liang, Jiuguang Wang, Tomás Lozano-Pérez, Leslie Pack Kaelbling

机构 * MIT(麻省理工学院) Princeton University(普林斯顿大学) University of Cambridge(剑桥大学) RAI Institute(RAI研究院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 通过预训练视觉-语言模型学习符号世界模型,以实现复杂机器人领域中长周期决策制定的零样本泛化。

Comments A version of this paper appears in the official proceedings of RA-L, Volume 11, Issue 4

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07545 2026-03-10 cs.CV cs.AI cs.LG 94%

DreamSAC: Learning Hamiltonian World Models via Symmetry Exploration

DreamSAC:通过对称探索学习哈密顿世界模型

Jinzhou Tang, Fan Feng, Minghao Fu, Wenjun Lin, Biwei Huang, Keze Wang

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 DreamSAC通过基于哈密顿的对称探索方法,学习物理不变性以提升3D物理模拟中的外推泛化能力。

Comments 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07799 2026-03-10 cs.CV cs.RO 94%

MWM: Mobile World Models for Action-Conditioned Consistent Prediction

MWM: 移动世界模型用于动作条件一致预测

Han Yan, Zishang Xiang, Zeyu Zhang, Hao Tang

机构 * School of Computer Science, Peking University(北京大学计算机科学系)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 MWM提出了一种移动世界模型,通过两阶段训练框架和推理一致状态蒸馏提升动作条件一致性的图像目标导航性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14357 2026-03-10 cs.CV cs.LG 94%

Vid2World: Crafting Video Diffusion Models to Interactive World Models

Vid2World: 构建视频扩散模型以交互式世界模型

Siqiao Huang, Jialong Wu, Qixing Zhou, Shangchen Miao, Mingsheng Long

机构 * Tsinghua University(清华大学) Chongqing University(重庆大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 Vid2World通过因果化视频扩散模型,提升其在交互式世界建模中的可控性和生成能力,适用于机器人操作、游戏模拟和开放世界导航等多种领域。

Comments Project page: http://knightnemo.github.io/vid2world/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06987 2026-03-10 cs.RO cs.AI 93%

Foundational World Models Accurately Detect Bimanual Manipulator Failures

基础世界模型准确检测双臂机械臂故障

Isaac R. Ward, Michelle Ho, Houjun Liu, Aaron Feldman, Joseph Vincent, Liam Kruse, Sean Cheong, Duncan Eddy, Mykel J. Kochenderfer, Mac Schwager

机构 * Stanford University(斯坦福大学) Watney Robotics(Watney机器人公司)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文提出基于视觉基础模型的双臂机械臂故障检测方法,通过压缩潜在空间中的世界模型提升检测精度,相比传统方法在参数效率和故障检测率上均表现更优。

Comments 8 pages, 5 figures, accepted at the 2026 IEEE International Conference on Robotics and Automation

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08519 2026-03-10 cs.RO 92%

AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models

AtomVLA: 通过预测性潜在世界模型实现机器人操作的可扩展后训练

Xiaoquan Sun, Zetian Xu, Chen Cao, Zonghe Liu, Yihan Sun, Jingrui Pang, Ruijian Zhang, Zhen Yang, Kang Pang, Dingxin He, Mingqi Yuan, Jiayu Chen

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 AtomVLA通过预测性潜在世界模型实现机器人操作的可扩展后训练,提升长时间任务的鲁棒性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07562 2026-03-10 cs.CV 90%

Brain-WM: Brain Glioblastoma World Model

Brain-WM: 脑部胶质瘤世界模型

Chenhui Wang, Boyun Zheng, Liuxin Bao, Zhihao Peng, Peter Y. M. Woo, Hongming Shan, Yixuan Yuan

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 Brain-WM通过统一治疗预测与MRI生成,实现了肿瘤与治疗的共进化动态建模,提升了治疗计划的准确率和MRI生成的质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10480 2026-03-10 cs.CL 90%

Neuro-Symbolic Synergy for Interactive World Modeling

神经符号协同用于交互世界建模

Hongyu Zhao, Siyu Zhou, Haolin Yang, Zengyi Qin, Tianyi Zhou

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 NeSyS通过结合LLMs的概率语义先验与可执行符号规则,实现交互世界建模的表达力与鲁棒性,提升预测准确性和数据效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11892 2026-03-10 cs.CL 90%

R-WoM: Retrieval-augmented World Model For Computer-use Agents

R-WoM:基于检索的世界模型用于计算机使用代理

Kai Mei, Jiang Guo, Shuaichen Chang, Mingwen Dong, Dongkyu Lee, Xing Niu, Jiarong Jiang

机构 * Rutgers University(罗杰斯大学) AWS Agentic AI(亚马逊敏捷人工智能)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 R-WoM通过整合外部检索知识提升世界模型能力,有效改善长周期模拟中的决策表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07039 2026-03-10 cs.AI 89%

Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

具有4D空间-时间嵌入的自监督多模态世界模型

Lance Legel, Qin Huang, Brandon Voelker, Daniel Neamati, Patrick Alan Johnson, Favyen Bastani, Jeff Rose, James Ryan Hennessy, Robert Guralnick, Douglas Soltis, Pamela Soltis, Shaowen Wang

机构 * Ecological Intelligence Lab(生态智能实验室) School of Complex Adaptive Systems(复杂适应系统学院) University of Houston(休斯顿大学) Geosensing Systems Engineering & Sciences Lab(传感系统工程与科学实验室) Stanford University(斯坦福大学) Allen Institute for Artificial Intelligence(人工智能研究院) Spatial Intelligence Lab(空间智能实验室) Department of Computer Science(计算机科学系) Georgia Institute of Technology(佐治亚理工学院) Florida Museum of Natural History(佛罗里达自然历史博物馆) University of Florida(佛罗里达大学) NSF Institute for Geospatial Understanding(国家科学基金会地理理解研究所) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 DeepEarth通过4D空间-时间嵌入实现自监督多模态世界模型,在生态预测中取得最佳性能。

Comments 8 pages, 5 figures, 1 table. Presented at 2026 World Modeling Workshop, Mila Quebec

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08455 2026-03-10 cs.AI cs.LG 88%

The Boiling Frog Threshold: Criticality and Blindness in World Model-Based Anomaly Detection Under Gradual Drift

青蛙沸腾阈值:世界模型基于异常检测中的临界性与盲目性

Zhe Hong

机构 * National University of Singapore(新加坡国立大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI、cs.LG

AI总结 研究揭示了世界模型在异常检测中的临界阈值,发现检测阈值受噪声底座、检测器和环境动态的三重交互影响,且正弦漂移无法被检测到。

Comments 10 pages, 5 figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11682 2026-03-10 cs.RO cs.AI cs.SY eess.SY 88%

Ego-Vision World Model for Humanoid Contact Planning

人形机器人接触规划的视角世界模型

Hang Liu, Yuman Gao, Sangli Teng, Yufeng Chi, Yakun Sophia Shao, Zhongyu Li, Maani Ghaffari, Koushil Sreenath

机构 * University of California, Berkeley(加州大学伯克利分校) University of Michigan, Ann Arbor(密歇根大学安娜堡分校) The Chinese University of Hong Kong(香港中文大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI、cs.RO

AI总结 本文提出了一种结合学习世界模型与MPC的框架,用于提升人形机器人在复杂环境中的接触规划能力,实现更高效和稳健的多任务执行。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23488 2026-03-10 cs.AI cs.CL 69%

Mapping Overlaps in Benchmarks through Perplexity in the Wild

通过在野 perplexity 映射重叠关系

Siyang Wu, Honglin Bao, Sida Li, Ari Holtzman, James A. Evans

机构 * Data Science Institute, University of Chicago(芝加哥大学数据科学研究所)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 通过分析LLM基准测试的perplexity,揭示了不同任务间的重叠结构及LLM能力差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08024 2026-03-10 cs.CL 67%

ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments

ConflictBench: 通过交互式和视觉 grounded 环境评估人类-人工智能冲突

Weixiang Zhao, Haozhen Li, Yanyan Zhao, xuda zhi, Yongbo Huang, Hao He, Bing Qin, Ting Liu

机构 * Harbin Institute of Technology(哈尔滨工业大学) SERES

专题命中 通用世界模型 :world model(abstract);world model(abstract)

AI总结 ConflictBench通过交互式和视觉 grounded 环境评估人类-人工智能冲突,揭示智能体在不同风险情境下的行为差异及对齐失败问题。

Comments 29 pages, 20 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07409 2026-03-10 stat.ME stat.ML 64%

Tree-Based Predictive Models for Noisy Input Data

基于树模型的噪声输入数据预测方法

Kevin McCoy, Zachary Wooten, Christine B. Peterson

专题命中 通用世界模型 :predictive model(title,abstract);predictive models(title,abstract)

AI总结 本文提出meBART模型,通过直接整合测量误差来提升预测精度和不确定性量化能力,适用于生物医学领域中存在测量误差的预测问题。

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12945 2026-03-10 cs.CL cs.LG 50%

A Component-Based Survey of Interactions between Large Language Models and Multi-Armed Bandits

基于组件的大型语言模型与多臂老虎机交互调研

Siguang Chen, Chunli Lv, Miao Xie

机构 * College of Information and Electrical Engineering, China Agricultural University, Beijing 100083, China(信息与电气工程学院,中国农业大学,北京) Key Laboratory of Agricultural Machinery Monitoring and Big Data Application, Ministry of Agriculture and Rural Affairs, Beijing 100083, China(农业机械监测与大数据应用重点实验室,农业农村部,北京)

专题命中 通用世界模型 :environment model(abstract);分类 cs.LG

AI总结 本文调研了大型语言模型与多臂老虎机在组件层面的双向交互,分析了两者在解决对方挑战中的协同作用及未来研究方向。

Comments 25 pages, 6 table

详情

展开后加载摘要…

URL PDF HTML 收藏