arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6483 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4335 篇

1803.01901 2018-03-07 cs.LG cs.AI cs.CY stat.ML 50%

On Discrimination Discovery and Removal in Ranked Data using Causal Graph

Yongkai Wu, Lu Zhang, Xintao Wu

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.10651 2018-02-02 stat.ML cs.AI cs.LG 50%

Reliable Decision Support using Counterfactual Models

Peter Schulam, Suchi Saria

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

Comments Published in the proceedings of Neural Information Processing Systems (NIPS) 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.06136 2017-11-17 cs.CV 50%

3D Trajectory Reconstruction of Dynamic Objects Using Planarity Constraints

Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen

专题命中 通用世界模型 :environment model(abstract);分类 cs.CV

Comments 9 Pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.02604 2017-02-24 cs.LG cs.AI cs.NE stat.ML 50%

Causal Regularization

Mohammad Taha Bahadori, Krzysztof Chalupka, Edward Choi, Robert Chen, Walter F. Stewart, Jimeng Sun

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

Comments Adding theoretical analysis, revising the text

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.02184 2017-02-09 cs.LG 50%

Transfer from Multiple Linear Predictive State Representations (PSR)

Sri Ramana Sekharan, Ramkumar Natarajan, Siddharthan Rajasekaran

专题命中 通用世界模型 :environment model(abstract);分类 cs.LG

Comments 8 pages, 3 algorithms, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.04977 2016-12-16 cs.SE cs.CY cs.RO 50%

Towards the Verification of Safety-critical Autonomous Systems in Dynamic Environments

Adina Aniculaesei, Daniel Arnsberger, Falk Howar, Andreas Rausch

专题命中 通用世界模型 :environment model(abstract);分类 cs.RO

Comments In Proceedings V2CPS-16, arXiv:1612.04023

Journal ref EPTCS 232, 2016, pp. 79-90

详情

展开后加载摘要…

URL PDF HTML 收藏
1604.02855 2016-10-06 stat.ML cs.CV cs.LG 50%

Active Learning for Online Recognition of Human Activities from Streaming Videos

Rocco De Rosa, Ilaria Gori, Fabio Cuzzolin, Barbara Caputo, Nicolò Cesa-Bianchi

专题命中 通用世界模型 :分类 cs.LG、cs.CV;predictive model(abstract);predictive models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.04906 2016-07-15 cs.LG cs.CV 50%

Performing Highly Accurate Predictions Through Convolutional Networks for Actual Telecommunication Challenges

Jaime Zaratiegui, Ana Montoro, Federico Castanedo

专题命中 通用世界模型 :分类 cs.LG、cs.CV;predictive model(abstract);predictive models(abstract)

Comments 11 pages, 6 figures, accepted by IJCAI-16 Workshop on Deep Learning for Artificial Intelligence (DLAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.04126 2016-01-19 stat.ML cs.AI cs.CY cs.LG 50%

Engineering Safety in Machine Learning

Kush R. Varshney

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

Comments 2016 Information Theory and Applications Workshop, La Jolla, California

详情

展开后加载摘要…

URL PDF HTML 收藏
1301.0600 2015-05-19 cs.LG cs.AI cs.IR 50%

An MDP-based Recommender System

Guy Shani, Ronen I. Brafman, David Heckerman

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

Comments Appears in Proceedings of the Eighteenth Conference on Uncertainty in Artificial Intelligence (UAI2002)

详情

展开后加载摘要…

URL PDF HTML 收藏
1201.0979 2015-03-19 cs.LO cs.AI cs.PL 50%

Sciduction: Combining Induction, Deduction, and Structure for Verification and Synthesis

Sanjit A. Seshia

专题命中 通用世界模型 :environment model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1502.00062 2015-02-03 stat.ML cs.AI cs.LG 50%

A New Intelligence Based Approach for Computer-Aided Diagnosis of Dengue Fever

Vadrevu Sree Hari Rao, Mallenahalli Naresh Kumar

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

Comments 7 pages, 5 figures. arXiv admin note: substantial text overlap with arXiv:1501.07093

Journal ref Information Technology in Biomedicine, IEEE Transactions on , vol.16, no.1, pp.112,118, Jan. 2012

详情

展开后加载摘要…

URL PDF HTML 收藏
1408.6127 2014-08-27 cs.AI 50%

A Complete framework for ambush avoidance in realistic environments

Emmanuel Boidot, Aude Marzuoli, Eric Feron

专题命中 通用世界模型 :environment model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1306.1553 2013-06-26 cs.AI 50%

Direct Uncertainty Estimation in Reinforcement Learning

Sergey Rodionov, Alexey Potapov, Yurii Vinogradov

专题命中 通用世界模型 :environment model(abstract);分类 cs.AI

Comments AGI-13 Workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1111.3934 2012-05-15 cs.AI 50%

Model-based Utility Functions

Bill Hibbard

专题命中 通用世界模型 :environment model(abstract);分类 cs.AI

Comments 24 pages, extensive revisions

Journal ref Journal of Artificial General Intelligence 3(1) 1-24, 2012

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频世界模型 53 篇

2608.09926 2026-08-11 cs.CV 新提交 96%

Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

学习世界如何演化:基于潜动态推理的外推视频世界模型

Haodong Li, Shaoteng Liu, Tianyu Wang, Chongjian Ge, Sihui Ji, Jiahan Zhang, Xin Lin, Haolin Lu, Zhe Lin, Manmohan Chandraker

机构 * UCSD(加利福尼亚大学圣迭戈分校) Adobe(奥多比公司)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);latent dynamics(title,abstract);world models(title)

AI总结 针对主流视频扩散模型未建模像素时间转换的问题,提出LDR方法,在PhyWorld基准上实现更优动态外推,参数更少、速度更快,是首个能泛化至训练分布外的视频世界模型

Comments Project page: https://lat-dyn-reason.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31158 2026-06-19 cs.CV cs.LG 版本更新 96%

Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models

光交互:交互式视频世界模型的免训练推理加速

Jiacheng Lu, Haoyi Zhu, Sipei Yi, Enze Xie, Yu Li, Cheng Zhuo

机构 * Zhejiang University(浙江大学) NVIDIA

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 针对交互式视频世界模型推理成本高的问题,提出免训练加速框架Light Interaction,通过自适应上下文管理、去噪缓存加速和3D块稀疏注意力实现最高2.59倍加速。

Comments 13 pages, 6 figures, 3 tables. Project page: https://2843721358l-del.github.io/Light-Interaction-Project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22882 2026-08-05 cs.CV cs.RO 版本更新 96%

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation

GEM-4D:用于机器人操作的几何增强视频世界模型

Kaichen Zhou, Yuzhen Chen, Fangneng Zhan, Hang Hua, Grace Chen, Xinhai Chang, Ao Qu, Yilun Du, Zhuang Liu, Paul Pu Liang, Mengyu Wang

机构 * Harvard AI and Robotics Lab(哈佛人工智能与机器人实验室) Harvard University(哈佛大学) Media Lab and EECS(媒体实验室和电子工程与计算机科学系) MIT(麻省理工学院) Princeton University(普林斯顿大学) MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出GEM-4D,通过注入从预训练几何基础模型蒸馏的密集4D对应监督,增强视频世界模型的几何一致性,并引入逆动力学模块将视频滚动转换为可执行机器人轨迹,提升操作成功率。

Comments Robotic World Model, Video Generative Model

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00267 2026-06-02 cs.CV cs.AI cs.LG cs.RO 96%

StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement

StressDream: 引导视频世界模型实现鲁棒的策略评估与改进

Junwon Seo, Sushant Veer, Ran Tian, Wenhao Ding, Apoorva Sharma, Karen Leung, Edward Schmerling, Marco Pavone, Andrea Bajcsy

机构 * Carnegie Mellon University(卡内基梅隆大学) NVIDIA Research(NVIDIA研究) University of Washington(华盛顿大学) Stanford University(斯坦福大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出StressDream方法,通过优化扩散视频世界模型的初始噪声,在推理时引导生成高影响且合理的未来场景,以支持鲁棒的策略评估与改进。

Comments Project page: https://junwon.me/StressDream/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05948 2026-08-07 cs.AI cs.CV cs.RO 新提交 96%

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

GAUGE:面向仿真引擎与视频世界模型物理保真度的、基于测量的基准

Shuai Wang, Yaxin Feng, Xuekun Jiang, Shihan Tian, Ningyu Yan, Xing Shen, Chaoyang Lyu, Hui Wang, Yunsong Zhou, Hanqing Wang, Jiangmiao Pang, Yang Xiang, Xing Gao, Chunhua Shen, Weinan Zhang

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 本研究提出基于真实世界的GAUGE基准,联合评估数值仿真器与生成式视频世界模型的物理保真度,发现无统一保真的物理引擎,视频模型存在轨迹形式正确但物理量错误的问题,为开发高保真仿真器奠定基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28362 2026-07-31 cs.CV cs.AI cs.LG 新提交 96%

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

ShadowDancer:通过从视频及其阴影中学习统一动力学表示,赋予视频世界模型任意动作控制能力

Jin Cao, Zian Meng, Kaipeng Zhang

机构 * Alaya Lab(Alaya实验室) Shanghai Innovation Institute(上海创新研究院)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 ShadowDancer通过阴影对与跨阴影预测学习统一动力学表示,实现视频世界模型的任意动作帧级控制,在多动力学族上的动作迁移与长回放性能优于基线,平均盲胜率达86%。

Comments https://ShadowDancer-1.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17808 2026-03-25 cs.RO cs.AI 96%

EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards

EVA:通过逆动力学奖励对齐视频世界模型与可执行机器人动作

Ruixiang Wang, Qingming Liu, Yueci Deng, Guiliang Liu, Zhen Liu, Kui Jia

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) DexForce Technology Co., Ltd.(DexForce技术有限公司)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 本文提出EVA框架,通过逆动力学奖励对齐视频世界模型与可执行动作,减少生成动作中的体感约束违规,提升下游任务执行成功率。

Comments Project page: https://eva-project-page.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16829 2026-08-18 cs.LG cs.AI 新提交 96%

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

CaliBench:视频世界模型的随机动力学是否经过物理校准?

Jonathan Sadeghi, Jenny Seidenschwarz, Jesse Allardice, Sirish Srinivasan, Benjamin Graham, Jeffrey Hawke

机构 * Odyssey(奥德赛公司)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 CaliBench用于测试视频世界模型的物理校准,将性能分解为可评分性与校准度,发现多数场景-模型组合显著未校准,发布了mnTV指标用于模型对比。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07408 2026-08-10 cs.CV cs.LG 新提交 96%

Addressable Memory for Video World Models

视频世界模型的可寻址内存

Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taixé, Despoina Paschalidou, Jonathan Lorraine, Aljoša Ošep

机构 * NVIDIA(英伟达) Princeton University(普林斯顿大学) University of Toronto(多伦多大学) Vector Institute(矢量研究院)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 针对交互式视频世界模型超出训练时序后内存寻址失效及压缩缓存破坏内存的问题,提出无训练框架 WorldTrace,含两种压缩方法,在新基准 LoopBench 上分别提升时序一致性 15.5%、情景回忆 19.5%。

Comments Project page: https://research.nvidia.com/labs/sil/projects/WorldTrace/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04964 2026-08-06 cs.AI cs.LG 新提交 96%

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

WorldCycle:用于长时序视频世界模型的自验证强化学习

Bohai Gu, Yueyang Yuan, Taiyi Wu, Dazhao Du, Jian Liu, Xiaoyi Pang, Jie Zhang, Xiaocheng Lu, Haobin Zhong, Xiaotong Zhao, Alan Zhao, Song Guo

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 本研究针对交互式视频世界模型的误差累积问题,提出自验证RL框架WorldCycle,利用可逆动作循环实现无标注监督,在CycleBench基准上大幅降低状态返回漂移并提升复合动作准确率。

Comments https://nevsnev.github.io/Worldcycle/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01600 2026-06-02 cs.CV cs.CL cs.RO 96%

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

RoboTrustBench:机器人操作视频世界模型的可信度基准测试

Huiqiong Li, Jiayu Wang, Zhiting Mei, Anirudha Majumdar, Jingjing Chen, Bin Zhu

机构 * Singapore Management University(新加坡国立管理学院) Fudan University(复旦大学) Princeton University(普林斯顿大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 针对视频世界模型在机器人操作中的可信度问题,提出RoboTrustBench基准,包含正常、约束敏感、反事实和对抗四种场景,通过专家验证的指令-图像对和六维评估协议,发现当前模型在约束推理、反事实基础、物理交互和不安全指令抑制方面存在不足。

Comments Project: https://huiqiongli.github.io/RoboTrustBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17792 2026-04-16 cs.CV cs.RO 96%

Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?

Target-Bench: 视频世界模型能否在语义目标下实现无地图路径规划?

Dingrui Wang, Zhihao Liang, Hongyuan Ye, Zhexiao Sun, Zhaowei Lu, Yuchen Zhang, Yuyu Zhao, Yuan Gao, Marvin Seegert, Finn Schäfer, Haotong Qin, Wei Li, Luigi Palmieri, Felix Jahncke, Mattia Piccinini, Johannes Betz

机构 * TUM(慕尼黑技术大学) Bosch AI Center(博世人工智能中心) ETH(苏黎世联邦理工学院) NJU(南京大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 Target-Bench首次评估视频世界模型的语义推理、空间估计和规划能力,通过450个机器人采集场景和5个互补指标揭示当前模型在语义推理与视觉生成间的显著差距。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25716 2026-03-31 cs.CV cs.AI 96%

Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

视线之外但未被遗忘:动态视频世界模型的混合记忆

Kaijin Chen, Dingkang Liang, Xin Zhou, Yikang Ding, Xiaoqiang Liu, Pengfei Wan, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Kling Team, Kuaishou Technology(快手科技Kling团队)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 本文提出混合记忆机制,解决动态主体消失后重现时的连续性问题,通过HM-World数据集和HyDRA架构提升视频世界模型的动态一致性与生成质量。

Comments Project Page: https://kj-chen666.github.io/Hybrid-Memory-in-Video-World-Models/ Code: https://github.com/H-EmbodVis/HyDRA

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15043 2026-08-18 cs.AI 新提交 96%

SCOPE: Score-Isolated Agentic Optimization for Video World Models

SCOPE:用于视频世界模型的分数隔离智能体优化

Yuhua Jiang, Jiaming Wang, Qingbin Liu, Feifei Gao

机构 * Tsinghua University(清华大学) National University of Singapore(新加坡国立大学) Tencent(腾讯)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 本研究针对视频世界模型推理时改进的评估难题,提出SCOPE框架,通过类型化状态更新与冻结策略实现可审计适应,在Physics-IQ基准上较冻结基线提升+14.24,同时揭示推理时更新的益处需原则性部署机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08982 2026-08-11 cs.LG 新提交 96%

Twin Rollouts: Noise-Coupled Counterfactual Branching in Interactive Video World Models

双回滚:交互式视频世界模型中的噪声耦合反事实分支

Yu Ma, Hongli Shi, Xinran Xu

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 本研究提出噪声耦合双回滚框架,解决交互式视频世界模型的反事实生成问题,规避近似逆问题,定义时空局部性度量,待开展实验验证。

详情

展开后加载摘要…

URL PDF HTML 收藏