arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-03-24 至 2026-03-24 共收录 18 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 18 篇

2603.22286 2026-03-24 cs.CV cs.AI cs.CL cs.LG 96%

WorldCache: Content-Aware Caching for Accelerated Video World Models

WorldCache: 用于加速视频世界模型的内容感知缓存

Umair Nawaz, Ahmed Heakl, Ufaq Khan, Abdelrahman Shaker, Salman Khan, Fahad Shahbaz Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence, UAE(穆罕默德·本·扎耶德人工智能大学,阿联酋)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 WorldCache通过引入运动自适应阈值、显著性加权漂移估计和最优近似方法,实现动态场景下的高效特征重用,提升推理速度并保持高质量。

Comments 33 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21546 2026-03-24 cs.LG cs.AI 94%

What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators

世界模型在强化学习中学习了什么?在学习的环境模拟器中探测潜在表示

Xinyu Zhang

机构 * Anyscale

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究通过线性与非线性探针、因果干预和注意力分析,探讨了IRIS和DIAMOND两种架构的世界模型对游戏状态的内部表示,发现其具有近似线性的结构化表示。

Comments 5 pages, 3 figures, 1 table

Journal ref ICLR 2026 the 2nd Workshop on World Models: Understanding, Modelling and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13340 2026-03-24 cs.RO cs.AI cs.LG 94%

Latent Policy Steering with Embodiment-Agnostic Pretrained World Models

基于世界模型的潜在策略引导

Yiqi Wang, Mrinal Verghese, Jeff Schneider

机构 * Robotics Institute, School of Computer Science, Carnegie Mellon University(机器人研究所、计算机科学学院、卡内基梅隆大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文通过预训练世界模型和优化价值函数,提升低数据场景下的视觉运动策略性能,采用光流作为动作表示,实现跨体素的策略优化,实验显示在机器人任务中性能提升显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21104 2026-03-24 cs.RO cs.CV 94%

CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation

CounterScene: 生成世界模型中的反事实因果推理用于安全关键闭环评估

Bowen Jing, Ruiyang Hao, Weitao Zhou, Haibao Yu

机构 * Tuojing Intelligence(途京智能) King's College London(伦敦大学国王学院) Tsinghua University(清华大学) The University of Hong Kong(香港大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 CounterScene通过结构化反事实推理生成安全关键驾驶场景,通过因果对抗代理识别和冲突感知交互世界模型,提升长周期碰撞率并保持轨迹真实性。

Comments 28 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22212 2026-03-24 cs.CV 93%

Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models

Omni-WorldBench:迈向全面的交互导向的世界模型评估

Meiqi Wu, Zhixin Cai, Fufangchen Zhao, Xiaokun Feng, Rujing Dang, Bingze Song, Ruitian Tian, Jiashu Zhu, Jiachen Lei, Hao Dou, Jing Tang, Lei Sun, Jiahong Wu, Xiangxiang Chu, Zeming Liu, Kaiqi Huang

机构 * School of Computer Science and Technology, UCAS(UCAS计算机科学与技术学院) The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, CASIA(复杂系统认知与决策智能重点实验室) School of Computer Science and Engineering, Beihang University(北航计算机科学与工程学院) State Key Laboratory of Networking and Switching Technology, BUPT(网络与交换技术国家重点实验室) AMAP, Alibaba Group(阿里妈妈实验室,阿里巴巴集团)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出Omni-WorldBench,一个针对4D世界模型交互响应能力的综合评估基准,通过Omni-WorldSuite和Omni-Metrics评估交互动作对状态转移的影响,分析现有模型的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21315 2026-03-24 cs.LG 93%

FluidWorld: Reaction-Diffusion Dynamics as a Predictive Substrate for World Models

FluidWorld: 反应-扩散动力学作为世界模型的预测性基质

Fabien Polly

机构 * Independent Researcher(独立研究者)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出FluidWorld,通过反应-扩散型偏微分方程实现世界模型预测,对比Transformer和ConvLSTM基线,展示PDE在参数效率和空间结构保留上的优势。

Comments 18 pages, 16 figures, 4 tables. Code available at https://github.com/infinition/FluidWorld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20607 2026-03-24 cs.RO cs.LG 93%

Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models

面向视觉-语言-动作模型的实用世界模型强化学习

Zhilong Zhang, Haoxiang Ren, Yihao Sun, Yifei Sheng, Haonan Wang, Haoxin Lin, Zhichao Wu, Pierre-Luc Bacon, Yang Yu

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China(新型软件技术国家实验室,南京大学,南京,中国) School of Artificial Intelligence, Nanjing University, Nanjing, China(人工智能学院,南京大学,南京,中国) Mila - Quebec AI Institute(魁北克AI研究所)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);model-based reinforcement learning(title);world models(abstract)

AI总结 本文提出VLA-MBPO框架,解决视觉-语言-动作模型在强化学习中的世界建模、多视角一致性及稀疏奖励下的误差累积问题,提升策略性能和样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21998 2026-03-24 cs.CV cs.RO 92%

Causal World Modeling for Robot Control

机器人控制中的因果世界建模

Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Fei Han, Mingrui Yu, Zelin Gao, Nan Xue, Xing Zhu, Yujun Shen, Yinghao Xu

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);video world model(abstract)

AI总结 本文提出LingBot-VA框架,通过视频世界建模与视觉语言预训练结合,实现机器人学习的新基础。模型包含共享潜在空间、闭环回滚机制和异步推理流程,提升了长周期操作和数据效率。

Comments Project page: https://technology.robbyant.com/lingbot-va Code: https://github.com/robbyant/lingbot-va

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21017 2026-03-24 cs.RO 89%

Dreaming the Unseen: World Model-regularized Diffusion Policy for Out-of-Distribution Robustness

梦见未见:用于分布外鲁棒性的世界模型正则化扩散策略

Ziou Hu, Xiangtong Yao, Yuan Meng, Zhenshan Bing, Alois Knoll

机构 * School of Computation, Information and Technology, Technical University of Munich, Garching, Germany(慕尼黑技术大学计算与信息学院) State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) the School of Science and Technology, Nanjing University (Suzhou Campus), China(南京大学苏州校区科学技术学院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);latent dynamics(abstract);分类 cs.RO

AI总结 本文提出Dream Diffusion Policy,通过整合世界模型提升扩散策略在分布外扰动下的鲁棒性,实验显示其在MetaWorld和现实场景中表现优异。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21557 2026-03-24 cs.CV 88%

From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy

从部分到整体:具有自适应结构层次的3D生成世界模型

Bi'an Du, Daizong Liu, Pufan Li, Wei Hu

机构 * Wangxuan Institute of \ Technology, Peking University Beijing, China Institute for Math \& AI, Wuhan University Wuhan, China

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.CV

AI总结 本文提出了一种自适应部分-整体层次的3D生成世界模型,通过直接从图像令牌推断出软且组合性的掩码,自主发现潜在结构槽,实现跨类别的形状共享和去噪。

Comments Accepted to ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21340 2026-03-24 cs.AI cs.DC 88%

ARYA: A Physics-Constrained Composable & Deterministic World Model Architecture

ARYA:一种受物理约束的可组合且确定性世界模型架构

Seth Dobrin, Lukasz Chmiel

机构 * ARYA Labs(ARYA实验室)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 本文提出ARYA,一种基于五项原则的可组合、受物理约束且确定性世界模型架构,通过层级系统实现高效能与计算效率的平衡,展示其在六个基准测试中的卓越表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20327 2026-03-24 cs.LG cs.AI cs.CV 87%

Probing the Latent World: Emergent Discrete Symbols and Physical Structure in Latent Representations

探测潜在世界:在潜在表示中涌现的离散符号与物理结构

Liu hung ming

机构 * PARRAWA AI(PARRAWA人工智能)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文提出AIM框架,通过被动量化探针探测V-JEPA 2潜在表示中的离散符号序列,揭示潜在空间的紧凑性及结构化符号 manifold 的可发现性。

Comments 35 pages, 6 figures, 3 tables, 26 equations; independent research report; Stage 1 of a four-stage AIM--V-JEPA 2 integration roadmap; code available at https://github.com/cyrilliu1974/JEPA

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05848 2026-03-24 cs.CV cs.AI cs.RO 83%

Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals

目标力:教视频模型实现物理条件化的目标

Nate Gillman, Yinghua Zhou, Zitian Tang, Evan Luo, Arjan Chakravarthy, Daksh Aggarwal, Michael Freeman, Charles Herrmann, Chen Sun

机构 * Brown University(布朗大学) Cornell University(康奈尔大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出Goal Force框架,通过显式力矢量和中间动力学定义目标,使视频模型能零样本泛化到复杂现实场景,实现基于物理的视频生成与规划。

Comments Camera ready version (CVPR 2026). Code and interactive demos at https://goal-force.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16177 2026-03-24 cs.LG 69%

The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data

微调者的谬误:何时应使用微调数据进行预训练

Christina Baek, Ricardo Pio Monti, David Schwab, Amro Abbas, Rishabh Adiga, Cody Blakeney, Maximilian Böther, Paul Burstein, Aldo Gael Carranza, Alvin Deng, Parth Doshi, Vineeth Dorna, Alex Fang, Tony Jiang, Siddharth Joshi, Brett W. Larsen, Jason Chan Lee, Katherine L. Mentzer, Luke Merrick, Haakon Mongstad, Fan Pan, Anshuman Suri, Darren Teh, Jason Telanoff, Jack Urbanek, Zhengping Wang, Josh Wills, Haoli Yin, Aditi Raghunathan, J. Zico Kolter, Bogdan Gaza, Ari Morcos, Matthew Leavitt, Pratyush Maini

机构 * DatologyAI Team(DatologyAI团队)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 本文研究了一种简单策略,即专门预训练(SPT),通过在预训练阶段重复使用小领域数据集,以提升领域性能并保留通用能力。实验显示,SPT能减少预训练token数量,提升领域表现,同时降低过拟合风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01641 2026-03-24 cs.CV 69%

FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion Deblurring

FideDiff:高效的高保真图像运动去模糊扩散模型

Xiaoyang Liu, Zhengyan Zhou, Zihang Xu, Jiezhang Cao, Zheng Chen, Yulun Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Harvard University(哈佛大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 本文提出FideDiff,一种高效的单步扩散模型,用于高保真图像运动去模糊。通过将运动去模糊转化为扩散过程,结合Kernel ControlNet和自适应时间步预测,提升了去模糊性能。

Comments Accepted to ICLR 2026. Code is available at https://github.com/xyLiu339/FideDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21778 2026-03-24 eess.SP cs.LG 58%

Cluster-Specific Predictive Modeling: A Scalable Solution for Resource-Constrained Wi-Fi Controllers

特定集群的预测建模:一种适用于资源受限Wi-Fi控制器的可扩展解决方案

Gianluca Fontanesi, Luca Barbieri, Lorenzo Galati Giordano, Alfonso Fernandez Duran, Thorsten Wild

机构 * Radio Systems Research, Nokia Bell Labs, Stuttgart, Germany(诺基亚贝尔实验室射频系统研究部,斯图加特,德国)

专题命中 通用世界模型 :predictive model(title,abstract);分类 cs.LG;predictive models(abstract)

AI总结 本文提出通过整合聚类算法和模型评估技术,优化大规模Wi-Fi网络中的预测建模,通过特征聚类提升特定集群的预测精度,展示集群特定模型在高活动集群中优于全局模型。

Comments 5 figures, 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22097 2026-03-24 cs.AI cs.LG 50%

SpecTM: Spectral Targeted Masking for Trustworthy Foundation Models

SpecTM:用于可信基础模型的频谱定向遮蔽

Syed Usama Imtiaz, Mitra Nasr Azadani, Nasrin Alamdari

机构 * Department of Civil and Environmental Engineering(土木与环境工程系)

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 本文提出SpecTM,一种结合物理约束的遮蔽设计,通过预训练时利用交叉频谱上下文重建目标波段,提升EO领域基础模型的可信度和可解释性。

Comments Accepted to IEEE IGARSS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06767 2026-03-24 cs.LG cs.AI 50%

Failure Detection in Chemical Processes Using Symbolic Machine Learning: A Case Study on Ethylene Oxidation

利用符号机器学习进行化学过程故障检测:乙烯氧化的案例研究

Julien Amblard, Niklas Groll, Matthew Tait, Mark Law, Gürkan Sin, Alessandra Russo

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 本文利用符号机器学习预测化学过程故障,通过乙烯氧化案例展示其在解释性与预测性能上的优势。

Comments Accepted at AAAI-MAKE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏