arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-04-08 至 2026-04-08 共收录 9 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 8 篇

2604.01346 2026-04-08 cs.CR cs.AI cs.LG cs.RO 94%

Safety, Security, and Cognitive Risks in World Models

世界模型中的安全性、安全性和认知风险

Manoj Parmar

机构 * SovereignAI Security Labs(SovereignAI安全实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文探讨了世界模型在自主决策中的安全、安全及认知风险,提出了轨迹持久性和表征风险的定义,并通过实验验证了对抗攻击的效果,强调了对世界模型的严谨性要求。

Comments version 2, 29 pages, 1 figure (6 panels), 3 tables. Empirical proof-of-concept on GRU/RSSM/DreamerV3 architectures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13009 2026-04-08 cs.CV 90%

Matrix-game 2.0: An open-source real-time and streaming interactive world model

矩阵游戏2.0:一个开源的实时和流式交互世界模型

Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, Baixin Xu, Hao-Xiang Guo, Kaixiong Gong, Size Wu, Wei Li, Xuchen Song, Yang Liu, Yangguang Li, Yahui Zhou

机构 * Skywork AI(天工AI)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出Matrix-Game 2.0,通过少步自回归扩散生成长视频,解决传统交互世界模型实时性差的问题,实现25FPS高速生成。

Comments Project Page: https://matrix-game-v2.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05930 2026-04-08 cs.CL cs.AI cs.LG cs.MA 90%

Can We Predict Before Executing Machine Learning Agents?

在执行机器学习代理之前,我们能否进行预测?

Jingsheng Zheng, Jintian Zhang, Yujie Luo, Yuren Mao, Yunjun Gao, Lun Du, Huajun Chen, Ningyu Zhang

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) Zhejiang University - Ant Group Joint Laboratory of Knowledge Graph(浙江大学-蚂蚁集团知识图谱联合实验室)

专题命中 通用世界模型 :world model(abstract,abstract_cn);world models(abstract,abstract_cn);world model(abstract,abstract_cn);world models(abstract,abstract_cn)

AI总结 本文提出通过内部化执行先验来预测数据驱动的解决方案偏好,利用预验证的数据分析报告提升LLM的预测能力,并在FOREAGENT中实现预测-验证循环,加速收敛并超越基于执行的基线。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05014 2026-04-08 cs.RO cs.AI cs.CV 87%

StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

StarVLA:一种积木式代码库,用于视觉-语言-动作模型开发

StarVLA Community

机构 * Von Neumann Institute, HKUST(香港科技大学冯·诺依曼研究所)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 StarVLA通过模块化架构、可重用训练策略和统一评估接口,解决VLA方法碎片化问题,提升可复现性和跨架构兼容性。

Comments Open-source VLA infra, Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11911 2026-04-08 cs.CV 81%

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model

InSpatio-WorldFM:一种开源的实时生成帧模型

InSpatio Team, Donghui Shen, Guofeng Zhang, Haomin Liu, Haoyu Ji, Jialin Liu, Jing Guo, Nan Wang, Siji Pan, Weihong Pan, Weijian Xie, Xiaojun Xiang, Xiaoyu Zhang, Xianbin Liu, Yifu Wang, Yipeng Chen, Zhewen Le, Zhichao Ye, Ziqiang Zhao

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出InSpatio-WorldFM,一种基于帧的实时生成模型,通过3D锚点和隐式空间记忆保持场景几何与细节,采用三阶段训练流程实现低延迟实时生成。

Comments Project page: https://inspatio.github.io/worldfm/ Code: https://github.com/inspatio/worldfm

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05015 2026-04-08 cs.CV 71%

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Video-MME-v2:迈向综合视频理解基准的下一阶段

Chaoyou Fu, Haozhi Yuan, Yuhao Dong, Yi-Fan Zhang, Yunhang Shen, Xiaoxing Hu, Xueying Li, Jinsen Su, Chengwu Long, Xiaoyao Xie, Yongkang Xie, Xiawu Zheng, Xue Yang, Haoyu Cao, Yunsheng Wu, Ziwei Liu, Xing Sun, Caifeng Shan, Ran He

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV;dynamics model(abstract)

AI总结 Video-MME-v2通过三级层次结构和非线性评估策略,评估视频理解的鲁棒性和准确性,揭示当前模型与人类专家间的差距及多模态推理瓶颈。

Comments Homepage: https://video-mme-v2.netlify.app/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18601 2026-04-08 cs.GR cs.AI cs.CV cs.LG 64%

BulletGen: Improving 4D Reconstruction with Bullet-Time Generation

BulletGen: 通过子弹时间生成提升4D重建

Denis Rozumny, Jonathon Luiten, Numair Khan, Johannes Schönberger, Peter Kontschieder

机构 * Meta Reality Labs(Meta 现实实验室)

专题命中 通用世界模型 :world model(comments);world models(comments);分类 cs.AI、cs.LG、cs.CV;world model(comments)

AI总结 BulletGen利用生成模型在Gaussian动态场景表示中校正错误并补全缺失信息,通过对扩散视频生成模型输出与4D重建在单个冻结子弹时间步骤的对齐,实现新颖视角合成和2D/3D跟踪任务的最先进结果。

Comments Accepted at CVPR 2026 Workshop "4D World Models: Bridging Generation and Reconstruction"

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05520 2026-04-08 eess.SP cs.AI 50%

Learned Elevation Models as a Lightweight Alternative to LiDAR for Radio Environment Map Estimation

学习的高程模型作为LiDAR的轻量级替代方案用于无线电环境图估计

Ljupcho Milosheski, Fedja Močnik, Mihael Mohorčič, Carolina Fortuna

机构 * Department of Communication Systems, Jožef Stefan Institute(Jožef Stefan研究所通信系统系)

专题命中 通用世界模型 :environment model(abstract);分类 cs.AI

AI总结 本文提出一种两阶段框架,利用卫星RGB图像学习高程模型,替代传统LiDAR数据,提升无线电环境图估计的精度和效率。

Comments 6 pages, 3 figures, 3 tables Submitted to PIMRC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 模型式强化学习 1 篇

2604.05185 2026-04-08 cs.LG cs.SY eess.SY 74%

Cross-fitted Proximal Learning for Model-Based Reinforcement Learning

基于交叉验证的近端学习用于基于模型的强化学习

Nishanth Venkatesh, Andreas A. Malikopoulos

机构 * Cornell University(康奈尔大学)

专题命中 模型式强化学习 :model-based reinforcement learning(title,abstract);分类 cs.LG

AI总结 本文提出一种交叉验证的近端学习方法,用于解决基于模型的强化学习中因隐藏混淆导致的模型偏差问题,通过更高效利用数据提高估计精度。

详情

展开后加载摘要…

URL PDF HTML 收藏