arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-08-28 至 2026-08-28 共收录 11 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 8 篇

2608.26190 2026-08-28 cs.AI 新提交 94%

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

用潜世界模型预测后果并强化导航策略

Zengmao Wang, Wei Gao, Shuhan Shen

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出用于机器人导航的兼容性预测潜世界模型(LWM),通过预测动作条件下的潜特征兼容性评估动作后果,可从无标注视频监督策略学习并经强化学习改进,在多机器人导航数据集上性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26239 2026-08-28 cs.RO 新提交 94%

WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression

WALL-SS:通过下一尺度自回归扩展来缩放长视界世界模型

Maeve Zhang, Rain Sun, Xiang Wang, Cyril Zhang, Shalfun Li, Meng Cao, Howard Lu, Ethan Chen, Harry Jhou, KZ Zheng, Lights Shi, Regis Cheng, Lorenzin, Robert Wang, Victor Yao, Gody Li, Elise Mon, Yohann Tang, Ryan Yu, PS Zhang, Vincent Chen, Hang Su, Roy Gan, Hao Wang, Qian Wang

机构 * X Square Robot(X Square机器人)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出WALL-SS世界模型,通过尺度自回归扩展实现动作可控的长视界机器人仿真,经实验验证其可提升动作跟随与轨迹精度,减少动作漂移和长视界不一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27367 2026-08-28 cs.CV cs.AI 新提交 93%

Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

连续容量增长:面向JEPA世界模型的视觉Transformer编码器的任务复杂度驱动的宽度与深度扩展

Frederik Berenz

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文提出SCG方法,使JEPA世界模型的视觉Transformer编码器可按需连续扩展宽度或深度,在多任务上提升性能并实现更高参数效率,且零误扩展。

Comments 12 pages, 2 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26214 2026-08-28 cs.CV 新提交 92%

Surgical Video Generation From Diffusion to World Models: A Survey

手术视频生成:从扩散模型到世界模型:综述

Fuxiang Huang, Chenxu Zhang, Liang Han, Lei Zhang

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 该综述梳理2024-2026年手术视频生成领域文献,分三类总结方法,指出生成任务从合成帧转向建模场景因果动态,分析瓶颈并提供实验参考,为相关交叉领域研究者提供参考。

Comments 4 pages, 1 figures, 3 tables. Accepted for oral presentation at the 2026 3rd International Conference on Intelligent Perception and Pattern Recognition (IPPR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27345 2026-08-28 cs.CV cs.AI 新提交 90%

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

PAWBench:我们离概率对齐的世界建模还有多远?

Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Krea AI Huggingface Shanghai Innovation Institute(上海创新研究院) Tongyi Lab(通义实验室) The University of Hong Kong(香港大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本研究针对当前视频生成器未满足概率对齐世界建模要求的问题,提出PAWBench基准与PAWEval协议,经50种场景和11个系统测试发现无模型达要求,为相关研究奠定基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27073 2026-08-28 cs.CV cs.RO 新提交 85%

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

SpatialCrafter:基于生成式3D代理的单图像世界建模

Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan

机构 * Hong Kong University of Science and Technology(香港科技大学) Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室) ManyCore Tech Inc.(ManyCore科技公司) Jilin University(吉林大学)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.CV、cs.RO

AI总结 SpatialCrafter是解决图像到场景生成问题的两阶段框架,通过3D代理及相关策略提升3D一致性,构建了115K场景的新数据集,性能优于现有方法且鲁棒性强。

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26200 2026-08-28 cs.AI cs.CV cs.LG 新提交 83%

GameWAM: A World Action Model for Video Games

GameWAM:面向电子游戏的世界动作模型

Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li

机构 * Fudan University(复旦大学) LIGHTSPEED(光速(企业名)) Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 GameWAM是首个用于原生闭环游戏玩法和GUI控制的世界动作模型,通过并行生成过程实现世界-动作联合学习,实验表明其以更少原生动作达到竞争力任务成功率,还发现了LASI失效模式。

Comments 44 pages, 23 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26151 2026-08-28 cs.AI cs.LG 新提交 50%

Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration

面向电信客户流失预测的可解释人工智能:一种CRM集成框架

Sandeep Gaddamwar

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 本文针对电信客户流失预测中模型不透明导致的CRM集成缺口,测试四种分类器并结合SHAP、LIME提供可解释性,提出四层CRM集成架构,预计可降低流失率3.3-5.3个百分点、节省19.9万-31.9万美元。

Comments 10 pages, 7 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频世界模型 2 篇

2608.27406 2026-08-28 cs.RO cs.AI cs.CV 新提交 96%

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

CLAP:跨具身视频世界模型是零样本物理模拟器

Kechen Liu, Ola Shorinwa

机构 * Princeton University(普林斯顿大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 CLAP是跨具身动作条件视频生成框架,可在人类与机器人的多样化视频上训练,能接近或超越单具身视频模型性能,开源代码与模型,为训练单具身视频世界模型提供新范式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27328 2026-08-28 cs.CV 新提交 96%

R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models

R2M-Bench:通过交互式视频世界模型中的相对一致性评估重访记忆

Qiwen Gu, Bingjie Gao, Rui Chen, Geng Li, Jifan Li, Qishuai Wen, Li Niu, Jing Tang, Xiangxiang Chu, Junqiao Zhao

机构 * DreamX Team, Alibaba Group(阿里巴巴集团DreamX团队) Tongji University(同济大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 该研究提出R2M-Bench基准,通过同滚动过程的相对校准评估视频世界模型的重访记忆,其NMR与人类判断相关性良好,可减少慢动作捷径,DreamX-World-Memo表现最优。

Comments Code: this https URL (https://github.com/AMAP-ML/R2MBench)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 仿真与规划 1 篇

2608.26788 2026-08-28 cs.AI cs.CL cs.MA cs.RO 新提交 73%

Decoupling Planning and Control for Instructable Agents

可指令智能体的规划与控制解耦

Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste, Ishita Dasgupta, Alane Suhr

机构 * UC Berkeley(加州大学伯克利分校) UBC(不列颠哥伦比亚大学) Google DeepMind(谷歌DeepMind)

专题命中 仿真与规划 :world-model(abstract);world-model(abstract);分类 cs.AI、cs.RO、cs.MA

AI总结 本研究提出Instruct-to-Act系统,将VLM规划器与世界模型控制器解耦,在七个具身环境中验证其性能优于仅控制器、直接VLM动作生成等基线,可替换VLM规划器且保持快速控制。

Comments Published as a conference paper at COLM 2026. Project page: this https URL (https://zinengtang.github.io/instruct-to-act/)

详情

展开后加载摘要…

URL PDF HTML 收藏