arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-08-28 至 2026-08-28 共收录 15 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 15 篇

2606.09828 2026-08-28 cs.CV 版本更新 96%

Latent Spatial Memory for Video World Models

视频世界模型的潜在空间记忆

Weijie Wang, Haoyu Zhao, Yifan Yang, Feng Chen, Zeyu Zhang, Yefei He, Zicheng Duan, Donny Y. Chen, Yuqing Yang, Bohan Zhuang

机构 * Zhejiang University(浙江大学) Microsoft Research(微软研究院) Adelaide University(阿德莱德大学) Monash University(莫纳什大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出潜在空间记忆框架Mirage,通过在扩散潜在空间中直接构建和查询3D缓存,避免像素空间重建,实现高效视频生成,速度提升10.57倍,内存减少55倍。

Comments Project Page: this https URL (https://aka.ms/latent-spatial-memory), Code: this https URL (https://github.com/microsoft/LatentSpatialMemory)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22294 2026-08-28 cs.RO 版本更新 94%

Beyond Instance Slots: Semantically Rich World Models for Physical Interaction Planning

超越实例槽:面向物理交互规划的语义丰富世界模型

Juntao Cheng, Jingkai Wang, Yijun Shen, Xiansheng Chen, Zhiwei Yu

机构 * Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究针对物理交互规划中实例槽无法明确实体任务角色的问题,提出SR-WM模型,通过功能角色绑定实现语义接口,在LIBERO模拟套件等评估中验证了其视觉动力学与规划决策的连接能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26190 2026-08-28 cs.AI 新提交 94%

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

用潜世界模型预测后果并强化导航策略

Zengmao Wang, Wei Gao, Shuhan Shen

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出用于机器人导航的兼容性预测潜世界模型(LWM),通过预测动作条件下的潜特征兼容性评估动作后果,可从无标注视频监督策略学习并经强化学习改进,在多机器人导航数据集上性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26239 2026-08-28 cs.RO 新提交 94%

WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression

WALL-SS:通过下一尺度自回归扩展来缩放长视界世界模型

Maeve Zhang, Rain Sun, Xiang Wang, Cyril Zhang, Shalfun Li, Meng Cao, Howard Lu, Ethan Chen, Harry Jhou, KZ Zheng, Lights Shi, Regis Cheng, Lorenzin, Robert Wang, Victor Yao, Gody Li, Elise Mon, Yohann Tang, Ryan Yu, PS Zhang, Vincent Chen, Hang Su, Roy Gan, Hao Wang, Qian Wang

机构 * X Square Robot(X Square机器人)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出WALL-SS世界模型,通过尺度自回归扩展实现动作可控的长视界机器人仿真,经实验验证其可提升动作跟随与轨迹精度,减少动作漂移和长视界不一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27367 2026-08-28 cs.CV cs.AI 新提交 93%

Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

连续容量增长:面向JEPA世界模型的视觉Transformer编码器的任务复杂度驱动的宽度与深度扩展

Frederik Berenz

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文提出SCG方法,使JEPA世界模型的视觉Transformer编码器可按需连续扩展宽度或深度,在多任务上提升性能并实现更高参数效率,且零误扩展。

Comments 12 pages, 2 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26214 2026-08-28 cs.CV 新提交 92%

Surgical Video Generation From Diffusion to World Models: A Survey

手术视频生成:从扩散模型到世界模型:综述

Fuxiang Huang, Chenxu Zhang, Liang Han, Lei Zhang

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 该综述梳理2024-2026年手术视频生成领域文献,分三类总结方法,指出生成任务从合成帧转向建模场景因果动态,分析瓶颈并提供实验参考,为相关交叉领域研究者提供参考。

Comments 4 pages, 1 figures, 3 tables. Accepted for oral presentation at the 2026 3rd International Conference on Intelligent Perception and Pattern Recognition (IPPR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15439 2026-08-28 cs.AI 版本更新 92%

Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?

解决ARC-AGI-3编码智能体是否需要可执行世界模型、简化和验证?

Sergey Rodionov

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 研究探讨解决ARC-AGI-3时编码智能体是否需可执行世界模型、简化和验证。通过四个基于Codex的嵌套智能体评估,发现各变体随模型和推理强度改进,组件影响因设置而异,完整验证处理最佳,文本变体在部分设置中表现出色。

Comments 45 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27345 2026-08-28 cs.CV cs.AI 新提交 90%

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

PAWBench:我们离概率对齐的世界建模还有多远?

Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Krea AI Huggingface Shanghai Innovation Institute(上海创新研究院) Tongyi Lab(通义实验室) The University of Hong Kong(香港大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本研究针对当前视频生成器未满足概率对齐世界建模要求的问题,提出PAWBench基准与PAWEval协议,经50种场景和11个系统测试发现无模型达要求,为相关研究奠定基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07775 2026-08-28 cs.CV 版本更新 88%

DALE-CT: Depth-Aware 2D Slice Encoders Learn an Anatomical World Model of Chest CT

DALE-CT: 用于计算机断层扫描的深度感知基础模型

Evan W. Damron, Mahmut S. Gokmen, Mitchell A. Klusty, Caroline N. Leach, Emily B. Collier, V. K. Cody Bumgardner

机构 * University of Kentucky(肯塔基大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.CV

AI总结 提出DALE-CT,一种基于LeJEPA的2D切片模型,通过3D深度感知预训练(利用解剖掩膜和异常标注)提升表示质量,在CT多异常检测中达到与3D视觉语言模型近似的性能。

Comments 18 pages, 5 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27073 2026-08-28 cs.CV cs.RO 新提交 85%

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

SpatialCrafter:基于生成式3D代理的单图像世界建模

Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan

机构 * Hong Kong University of Science and Technology(香港科技大学) Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室) ManyCore Tech Inc.(ManyCore科技公司) Jilin University(吉林大学)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.CV、cs.RO

AI总结 SpatialCrafter是解决图像到场景生成问题的两阶段框架,通过3D代理及相关策略提升3D一致性,构建了115K场景的新数据集,性能优于现有方法且鲁棒性强。

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26200 2026-08-28 cs.AI cs.CV cs.LG 新提交 83%

GameWAM: A World Action Model for Video Games

GameWAM:面向电子游戏的世界动作模型

Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li

机构 * Fudan University(复旦大学) LIGHTSPEED(光速(企业名)) Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 GameWAM是首个用于原生闭环游戏玩法和GUI控制的世界动作模型,通过并行生成过程实现世界-动作联合学习,实验表明其以更少原生动作达到竞争力任务成功率,还发现了LASI失效模式。

Comments 44 pages, 23 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25479 2026-08-28 cs.CV cs.AI 版本更新 82%

4DStreamCtrl: Interactive Video Generation with Online 4D Control

4DStreamCtrl:基于在线4D控制的交互式视频生成

Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu

机构 * Peking University(北京大学) Tencent Hunyuan(腾讯混元)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 4DStreamCtrl将相机运动、物体轨迹与深度统一为3D点轨迹表示,结合蒸馏的因果流式模型,首次实现交互式4D可控实时视频生成,精度与效率均优于现有方法。

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22642 2026-08-28 cs.LG cs.AI 版本更新 82%

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

Mol-JEPA:用于分子的多模态联合嵌入预测架构

Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff

机构 * University of Tübingen(蒂宾根大学) Boehringer Ingelheim(勃林格殷格翰) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Brown University(布朗大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对分子基础模型的化学无效增强等局限,提出Mol-JEPA多模态框架,利用模态掩码融入生化上下文,在基准测试中展现出优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03345 2026-08-28 cs.AI cs.CL 版本更新 81%

Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning

语言模型遵循奥卡姆剃刀吗?对归纳和反向推理中简约性的评估

Yunxin Sun, Abulhair Saparov

机构 * Purdue University(普渡大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文评估大型语言模型在归纳和反向推理中是否遵循奥卡姆剃刀原则,通过生成综合推理问题测试其在简单和复杂场景中的表现,发现模型在复杂世界模型中难以生成高质量假设。

Comments Accepted at EMNLP 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26151 2026-08-28 cs.AI cs.LG 新提交 50%

Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration

面向电信客户流失预测的可解释人工智能:一种CRM集成框架

Sandeep Gaddamwar

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 本文针对电信客户流失预测中模型不透明导致的CRM集成缺口,测试四种分类器并结合SHAP、LIME提供可解释性,提出四层CRM集成架构,预计可降低流失率3.3-5.3个百分点、节省19.9万-31.9万美元。

Comments 10 pages, 7 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏