arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2025-12-02 至 2025-12-02 共收录 16 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 12 篇

2511.19861 2025-12-02 cs.CV cs.RO 94%

GigaWorld-0: World Models as Data Engine to Empower Embodied AI

GigaWorld-0:世界模型作为数据引擎赋能具身AI

GigaWorld Team, Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Haoyun Li, Jiagang Zhu, Kerui Li, Mengyuan Xu, Qiuping Deng, Siting Wang, Wenkang Qin, Xinze Chen, Xiaofeng Wang, Yankai Wang, Yu Cao, Yifan Chang, Yuan Xu, Yun Ye, Yang Wang, Yukun Zhou, Zhengyuan Zhang, Zhehao Dong, Zheng Zhu

机构 * GigaWorld Team(GigaWorld团队)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 GigaWorld-0通过整合视频生成与3D建模技术,构建统一的数据引擎,生成高质量具身交互数据,提升VLA模型在现实任务中的性能。

Comments Project Page: https://giga-world-0.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07338 2025-12-02 cs.RO cs.AI 94%

Delta-Triplane Transformers as Occupancy World Models

Delta-Triplane Transformers 作为占用世界模型

Haoran Xu, Peixi Peng, Guang Tan, Yiqian Chang, Yisen Zhao, Yonghong Tian

机构 * School of Intelligent Systems Engineering, Shenzhen Campus of Sun Yat-sen University(南方科技大学深圳校区智能系统工程学院) Peng Cheng Laboratory(鹏城实验室) School of Electronic and Computer Engineering, Shenzhen Graduate School, Peking University(北京大学深圳研究生院电子与计算机工程学院) School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳校区计算机科学与技术学院) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 Delta-Triplane Transformers 通过紧凑的3D表示和增量预测策略,提升自动驾驶中占用世界模型的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01119 2025-12-02 cs.LG cs.AI 90%

World Model Robustness via Surprise Recognition

通过惊喜识别提升世界模型鲁棒性

Geigh Zollicoffer, Tanush Chopra, Mingkuan Yan, Xiaoxu Ma, Kenneth Eaton, Mark Riedl

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 通过惊喜识别提升世界模型在噪声环境下的鲁棒性,增强自动驾驶模拟中智能体的稳定性与性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17808 2025-12-02 cs.LG 90%

Remote Sensing-Oriented World Model

面向遥感的世界模型

Yuxi Lu, Biao Wu, Zhidong Li, Kunqi Li, Chenya Huang, Huacan Wang, Qizhen Lan, Ronghao Chen, Ling Chen, Bin Liang

机构 * University of Technology Sydney (UTS)(悉尼技术大学) University of Chinese Academy of Sciences (UCAS)(中国科学院大学) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) Peking University(北京大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出首个面向遥感的世界模型框架,通过RemoteBAGEL模型在RSWISE基准上实现空间推理的高效外推。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00005 2025-12-02 cs.RO 88%

DREAMer-VXS: A Latent World Model for Sample-Efficient AGV Exploration in Stochastic, Unobserved Environments

DREAMer-VXS:一种用于在随机、未观测环境中高效样本AGV探索的潜在世界模型

Agniprabha Chakraborty

机构 * Department of Power Engineering, Jadavpur University(功率工程系,贾瓦德普尔大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 DREAMer-VXS通过高效样本利用和潜在世界模型,实现AGV在随机未观测环境中的高效探索与鲁棒导航

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02016 2025-12-02 cs.CV 81%

Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Don't Know Galileo's Principle...for now

生成视频中的物体比看起来更慢:模型遭受次地球重力并目前还不知道伽利略原理...

Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad

机构 * Indian Institute of Science(印度科学研究院) Johns Hopkins University(约翰霍普金斯大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 研究发现视频生成器在重力表示上存在偏差,通过针对性适配可提升其重力模拟精度。

Comments https://gravity-eval.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01878 2025-12-02 cs.AI 75%

Graph Distance as Surprise: Free Energy Minimization in Knowledge Graph Reasoning

图距离作为惊喜:知识图谱推理中的自由能最小化

Gaganpreet Jhajj, Fuhua Lin

机构 * School of Computing(计算学院) Information Systems(信息系统) Athabasca University(亚伯达大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);model-based reinforcement learning(abstract);分类 cs.AI

AI总结 本文提出利用图距离最小化惊喜来改进知识图谱推理,通过连接自由能原理与KG系统,探索图距离在生成模型中的应用及其对语法结构的影响。

Comments Accepted to NORA Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01816 2025-12-02 cs.CV cs.AI 71%

Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights

Envision:基于因果世界过程洞察的统一理解和生成基准测试

Juanxi Tian, Siyuan Li, Conghui He, Lijun Wu, Cheng Tan

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI、cs.CV

AI总结 Envision通过因果事件进程基准测试,评估统一模型在动态时空一致性上的表现,揭示其在因果叙事一致性上的优势及仍需改进的时空一致性挑战。

Comments 35 pages, 12 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15605 2025-12-02 cs.RO cs.CL cs.CV 71%

SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models

SRPO:用于视觉-语言-动作模型的自指政策优化

Senyu Fei, Siyin Wang, Li Ji, Ao Li, Shiduo Zhang, Liming Liu, Jinlong Hou, Jingjing Gong, Xianzhong Zhao, Xipeng Qiu

机构 * Fudan University(复旦大学) Tongji University(同济大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV、cs.RO

AI总结 SRPO通过自指政策优化方法,利用模型自身生成的成功轨迹作为参考,有效解决VLA-RL中的奖励稀疏问题,实现高成功率和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01048 2025-12-02 cs.CV 69%

TRoVe: Discovering Error-Inducing Static Feature Biases in Temporal Vision-Language Models

TRoVe: 发现时间视觉-语言模型中的错误引发静态特征偏差

Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari, Curtis Langlotz

机构 * Stanford University(斯坦福大学) HOPPR

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 TRoVe通过识别时间视觉-语言模型中的错误引发静态特征偏差,提升模型在下游任务中的表现。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03938 2025-12-02 q-fin.MF q-fin.PM 64%

In-Sample and Out-of-Sample Sharpe Ratios for Linear Predictive Models

样本内与样本外的线性预测模型夏普比率

Antoine Jacquier, Johannes Muhle-Karbe, Joseph Mulligan

专题命中 通用世界模型 :predictive model(title,abstract);predictive models(title,abstract)

AI总结 本文研究了线性预测模型在样本外表现受过拟合影响的程度,通过计算夏普比率的闭式近似,发现复杂策略的样本外复制比率随训练数据量增加而提升。

Comments 36 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00289 2025-12-02 cs.LG cs.RO 50%

Data-Driven Modeling and Correction of Vehicle Dynamics

数据驱动建模与车辆动力学校正

Nguyen Ly, Caroline Tatsuoka, Jai Nagaraj, Jacob Levy, Fernando Palafox, David Fridovich-Keil, Hannah Lu

机构 * Department of Aerospace Engineering and Engineering Mechanics, The University of Texas at Austin(航空航天工程与工程力学系,德克萨斯大学奥斯汀分校) Department of Mathematics, The Ohio State University(数学系,俄亥俄州立大学) Department of Computer Science, The University of Texas at Austin(计算机科学系,德克萨斯大学奥斯汀分校) Oden Institute for Computational Engineering and Sciences, The University of Texas at Austin(计算工程与科学研究所,德克萨斯大学奥斯汀分校)

专题命中 通用世界模型 :分类 cs.LG、cs.RO;predictive model(abstract);predictive models(abstract)

AI总结 本文提出DRIPS和FML两种方法,用于数据驱动校正非自主车辆动力学,实现高效且准确的模型学习与误差修正。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身与机器人 1 篇

2512.01924 2025-12-02 cs.RO cs.AI cs.LG 89%

Real-World Robot Control by Deep Active Inference With a Temporally Hierarchical World Model

通过时序分层世界模型的深度主动推断实现现实世界的机器人控制

Kentaro Fujii, Shingo Murata

机构 * Graduate School of Integrated Design Engineering, Keio University(Keio大学整合设计工程研究院)

专题命中 具身与机器人 :world model(title,abstract);world model(title,abstract);分类 cs.AI、cs.LG、cs.RO

AI总结 本文提出一种结合时序分层世界模型的深度主动推断框架,用于在不确定环境中实现机器人高成功率的操作与探索性动作切换。

Comments Accepted for publication in IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 模型式强化学习 1 篇

2510.22969 2025-12-02 cs.AI cs.MA 78%

Multi-Agent Conditional Diffusion Model with Mean Field Communication as Wireless Resource Allocation Planner

多智能体条件扩散模型与均场通信作为无线资源分配规划器

Kechen Meng, Sinuo Zhang, Rongpeng Li, Xiangming Meng, Yansha Deng, Chan Wang, Ming Lei, Zhifeng Zhao

机构 * College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院) Zhejiang University-University of Illinois Urbana-Champaign (ZJU-UIUC) Institute, Zhejiang University(浙江大学-伊利诺伊大学厄巴纳-香槟分校联合研究所) Department of Engineering, King’s College London(伦敦国王学院工程系)

专题命中 模型式强化学习 :world model(abstract);world model(abstract);model-based RL(abstract);分类 cs.AI、cs.MA

AI总结 多智能体条件扩散模型与均场通信用于无线资源分配规划,通过扩散模型和逆动态模型提升样本效率与策略可扩展性,理论保证收敛稳定性,实验显示在无线网络优化中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 仿真与规划 2 篇

2506.06981 2025-12-02 cs.AI cs.LG 82%

Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended Environments

深度强化学习需要深度行为分析:通过无模型智能体在开放性环境中探索隐式规划

Riley Simmons-Edler, Ryan P. Badman, Felix Baastad Berg, Raymond Chua, John J. Vastola, Joshua Lunger, William Qian, Kanaka Rajan

机构 * Department of Neurobiology, Harvard Medical School(哈佛医学院神经生物学系) Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究学院) Department of Mathematics, NTNU(NTNU数学系) School of Computer Science, McGill University & Mila(麦吉尔大学计算机科学学院及Mila) Department of Computer Science, University of Toronto(多伦多大学计算机科学系) Biophysics Graduate Program, Harvard University(哈佛大学生物物理学研究生项目)

专题命中 仿真与规划 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文通过ForageWorld环境研究DRL智能体的行为,发现无模型智能体可通过涌现动态展现规划行为,提出通用分析框架用于研究复杂智能体的学习动态。

Comments Published at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00041 2025-12-02 cs.RO cs.CV 71%

VISTAv2: World Imagination for Indoor Vision-and-Language Navigation

VISTAv2:室内视觉与语言导航的世界想象

Yanjia Huang, Xianshun Jiang, Xiangbo Gao, Mingyang Wu, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯农工大学)

专题命中 仿真与规划 :world model(abstract);world model(abstract);分类 cs.CV、cs.RO

AI总结 VISTAv2通过生成式世界模型实现室内视觉与语言导航的稳健规划,结合动作条件想象和在线价值图融合,提升导航效率与鲁棒性。

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏