arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-06-05 至 2026-06-05 共收录 11 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 11 篇

2606.01935 2026-06-05 cs.CV 95%

Unified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and Planning

统一驾驶令牌:面向驾驶世界模型和规划的表示与几何引导的离散分词器

Ziyang Yao, Zeyu Zhu, YunCheng Jiang, Zibin Guo, Huijing Zhao

机构 * Peking University(北京大学) Xiaomi EV(小米电动车)

专题命中 通用世界模型 :world model(title,abstract);world models(title);driving world model(title);world model(title,abstract)

AI总结 提出一种表示引导与几何增强的离散分词器,通过联合监督学习紧凑令牌,同时优化重建保真度、表示一致性和规划性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05925 2026-06-05 cs.AI 93%

Towards World Models in Biomedical Research

迈向生物医学研究的世界模型

Guangyu Wang, Jingkun Yue, Siqi Zhang, Yu Liu, Xiaoyu Wang, Mingyuan Meng, Changwei Ji, Zongbo Han, Yulin Wang, Yang Yue, Frank Fu, Ting Chen, Song Wu, Ziwei Liu, Jiangning Song, Ming Li, Gao Huang, Xiaohong Liu, Athanasios Vasilakos, Xingcai Zhang, Ping Zhang, Yong Li

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China(网络与交换技术国家重点实验室,北京邮电大学,北京,中国) Department of Engineering Science, University of Oxford, Oxford, United Kingdom(英国牛津大学工程科学系,牛津,英国) Institute of Medical Artificial Intelligence, South China Hospital, Medical School, Shenzhen University, Shenzhen, Guangdong, China(医学人工智能研究所,南方医院,医学学院,深圳大学,深圳,广东,中国) Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China(中关村学院及中关村人工智能研究院,北京,中国) Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University, 100084, Beijing, China(北京信息科学与技术国家研究中心(BNRist),清华大学,100084,北京,中国) Department of Chemical and Nano Engineering, University of California, San Diego, La Jolla, CA, USA(美国加州大学圣地亚哥分校化学与纳米工程系,La Jolla,CA,美国) Nanyang Technological University, Singapore(新加坡南洋理工大学) Monash Biomedicine Discovery Institute and Department of Biochemistry and Molecular Biology, Monash University, Melbourne, Victoria, Australia(莫纳什大学生物医学发现研究所和生物化学与分子生物学系,墨尔本,维多利亚,澳大利亚) David R. Cheriton School of Computer Science, University of Waterloo, Waterloo, Ontario, Canada(加拿大滑铁卢大学戴维·R·切里顿计算机科学学校,滑铁卢,安大略,加拿大) Department of ICT and Center for AI Research, University of Agder (UiA), Jon Lilletuns vei 9, Grimstad, Norway(挪威阿格德大学(UiA)信息与通信技术系及人工智能研究中心,Jon Lilletuns vei 9,Grimstad,挪威) Department of Electronic Engineering, Tsinghua University, Beijing, China(清华大学电子工程系,北京,中国)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出生物医学世界模型作为AI驱动发现的新范式,通过学习分子、细胞、组织和临床状态的潜在表征及干预条件动态,实现未来轨迹模拟,并探讨其在虚拟细胞、类器官、虚拟患者和手术模拟等应用中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06147 2026-06-05 cs.AI 92%

WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation

WorldFly: 基于世界模型的视觉-语言-动作模型用于无人机导航

Shengtao Zheng, Kai Li, Weichen Zhang, Yu Meng, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang

机构 * Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) BNRist, Tsinghua University(清华大学北京研究院)

专题命中 通用世界模型 :world-model(title,abstract);world-model(title,abstract);world model(abstract);world models(abstract)

AI总结 提出WorldFly框架,通过双分支耦合流匹配机制联合生成未来视频预测和导航动作,解决城市峡谷中严重遮挡和视角剧变下的无人机导航问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05773 2026-06-05 cs.RO 92%

PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation

PiL-World: 用于VLA策略环内评估的块式世界模型

Chong Ma, Taiyi Su, Jian Zhu, Jianjun Zhang, Zitai Huang, Yi Xu, Hanli Wang

机构 * Tongji University(同济大学) AIRC, Midea Group(美的集团人工智能研究院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world-model(abstract)

AI总结 提出PiL-World,一种块式世界模型,通过交替VLA推理和世界模型预测实现闭环评估,无需真实机器人执行,显著降低成功率估计误差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04463 2026-06-05 cs.RO 92%

OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics

OSCAR: 面向机器人的全具身骨架条件世界动作模型

Zhuoyuan Wu, Jun Gao

机构 * Peking University(北京大学) University of Michigan(密歇根大学) NVIDIA(英伟达)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);video world model(abstract)

AI总结 提出OSCAR,一种基于动作条件的视频世界模型,通过大规模数据管道和2D骨架渲染统一表示,实现跨机器人具身的泛化,并用于策略评估。

Comments Project page: https://wuzy2115.github.io/oscar-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05979 2026-06-05 cs.RO cs.AI 88%

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

世界-语言-动作模型:统一世界建模、语言推理与动作合成

Yi Yang, Zhihong Liu, Siqi Kou, Yiyang Chen, Yanzhe Hu, Jianbo Zhou, Boyuan Zhao, Zhijie Wei, Xiao Xia, Xueqi Li, Pengfei Liu, Zhijie Deng

机构 * SJTU(上海交通大学) SII(上海研究院) HUST(华中科技大学) SCUT(华南理工大学) ECUST(东华大学) SHU(上海大学) NJUPT(南京工业大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI、cs.RO

AI总结 提出世界-语言-动作(WLA)模型,通过自回归Transformer联合预测文本子任务、子目标图像和机器人动作,融合世界建模与语言推理能力,实现多任务和长时域任务的最优性能。

Comments 19 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.03413 2026-06-05 cs.LG cs.AI 82%

Learning to Theorize the World from Observation

从观察中学习理论化世界

Doojin Baek, Gyubin Lee, Junyeob Baek, Hosung Lee, Sungjin Ahn

机构 * University of Washington(华盛顿大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 受认知科学启发,提出Learning-to-Theorize范式,通过神经理论家(NEO)模型从原始非文本观测中推断显式解释性理论,实现基于解释的泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19312 2026-06-05 cs.LG cs.AI 82%

LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

LeWorldModel:从像素稳定端到端联合嵌入预测架构

Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, Randall Balestriero

机构 * Mila & Université de Montréal(Mila与蒙特利尔大学) New York University(纽约大学) Samsung SAIL(三星SAIL) Brown University(布朗大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出LeWorldModel,一种通过仅使用两个损失项从原始像素稳定端到端训练的联合嵌入预测架构,显著减少了可调损失超参数,并在多种2D和3D控制任务中表现出色,同时在物理结构编码和物理不合理的事件检测方面展示了其能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05699 2026-06-05 cs.RO 81%

DexFuture: Hierarchical Future-State Visuomotor Targeting for Bimanual Dexterous Tool Use

DexFuture: 用于双手灵巧工具使用的分层未来状态视觉运动目标

Runfa Blark Li, Kuang-Ting Tu, Nikola Raicevic, Dwait Bhatt, Xinshuang Liu, Keito Suzuki, Ki Myung Brian Lee, Nikolay Atanasov, Truong Nguyen

机构 * UC San Diego(圣迭戈大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出DexFuture分层系统,通过高层未来状态视觉运动目标预测器和低层目标条件结构化灵巧策略,实现双手灵巧工具使用,达到90%的特权oracle性能,运行速度60Hz,比DexWM式CEM规划快约250倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05086 2026-06-05 hep-th 80%

Fermionic Kaluza-Klein mode mixing in braneworlds

膜世界中费米子Kaluza-Klein模式混合

Chun-E Fu, Wen-Xuan Ma

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 研究厚膜世界模型中背景扰动导致的费米子Kaluza-Klein模式混合,通过精确奇异值分解揭示宇称依赖的耦合与空间极化现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05555 2026-06-05 cs.LG cs.AI 76%

Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

表示学习实现可扩展的多任务深度强化学习

Johan Obando-Ceron, Lu Li, Scott Fujimoto, Pierre-Luc Bacon, Aaron Courville, Pablo Samuel Castro

机构 * Mila – Québec AI Institute(魁北克AI研究所) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) CIFAR AI Chair(CIFAR人工智能 chair) Google DeepMind(谷歌DeepMind)

专题命中 通用世界模型 :world-model(abstract);world-model(abstract);model-based RL(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种结合预测性表示学习与高容量值函数近似的无模型算法MR.Q,在无需规划的情况下,在多任务连续控制任务中超越基于世界模型的方法和多种深度强化学习基线,并显著降低计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏