Rishabh Sharma, Gijs Hogervorst, Wayne E. Mackey, David J. Heeger, Stefano Martiniani
机构
*
Simons Center for Computational Physical Chemistry(Simon计算物理化学中心)
;
Center for Soft Matter Research, Department of Physics(软物质研究中心,物理系)
;
Statespace Labs, Inc.(Statespace公司)
;
Center for Neural Science, New York University(神经科学中心,纽约大学)
;
Department of Psychology, New York University(心理学系,纽约大学)
;
Courant Institute of Mathematical Sciences, New York University(Courant数学科学研究所,纽约大学)
BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
BridgeV2W: 通过具身遮罩将视频生成模型与具身世界模型 bridging
Yixiang Chen, Peiyan Li, Jiabing Yang, Keji He, Xiangnan Wu, Yuan Xu, Kai Wang, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang
机构
*
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(模式识别新实验室(NLPR),自动化研究所,中国科学院)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Shandong University(山东大学)
CommentsThis paper has been accepted for presentation at the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (IEEE ICASSP 2026) Workshop: 'Multi-Modal Signal Processing and AI for Communications and Sensing in 6G and Beyond (MuSiC-6GB)'
Understanding World or Predicting Future? A Comprehensive Survey of World Models
理解世界还是预测未来?世界模型的全面综述
Jingtao Ding, Yunke Zhang, Yu Shang, Jie Feng, Yuheng Zhang, Zefang Zong, Yuan Yuan, Hongyuan Su, Nian Li, Jinghua Piao, Yucheng Deng, Nicholas Sukiennik, Chen Gao, Fengli Xu, Yong Li
机构
*
Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University(电子工程系,信息科学与技术国家研究中心(BNRist),清华大学)
机构
*
Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
;
State Key Laboratory of Multimedia Information Processing(国家多媒体信息处理重点实验室)
;
School of Computer Science, Peking University(北京大学计算机学院)
;
Hong Kong University of Science and Technology(香港科技大学)
机构
*
Hong Kong University of Science(香港科学与技术大学)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学)
GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
Anthony Chen, Wenzhao Zheng, Yida Wang, Xueyang Zhang, Kun Zhan, Peng Jia, Kurt Keutzer, Shanghang Zhang
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
Li Auto Inc.(利亚德公司)
;
UC Berkeley(加州大学伯克利分校)
WorldModelBench: Judging Video Generation Models As World Models
Dacheng Li, Yunhao Fang, Yukang Chen, Shuo Yang, Shiyi Cao, Justin Wong, Michael Luo, Xiaolong Wang, Hongxu Yin, Joseph E. Gonzalez, Ion Stoica, Song Han, Yao Lu