arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6483 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 视频世界模型 53 篇

2608.05070 2026-08-06 cs.CV 新提交 96%

HelloWorld: Enabling Socially Interactive Characters in Video World Models

HelloWorld:在视频世界模型中实现具有社交互动性的角色

Liangyang Ouyang, Ruicong Liu, Xuangeng Chu, Kaipeng Zhang, Yoichi Sato

机构 * The University of Tokyo(东京大学) Alaya Lab(阿莱亚实验室)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 该研究提出HelloWorld视频世界模型,通过自蒸馏流水线和训练模块实现角色与用户的社交互动,构建含400样本的基准,其互动质量优于基线且保持顶尖图像美学。

Comments Project page: https://github.com/AlayaLab/HelloWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01127 2026-08-05 cs.CV 版本更新 96%

MiniWorld: Democratizing the Training of Video World Models from Scratch

MiniWorld:从零开始普及视频世界模型的训练

Yian Zhao, Ruochong Zheng, Hongcan Guo, Yu Yan, Jian Zhang, Jie Chen

机构 * Peking University(北京大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 MiniWorld是从零开始训练流式视频世界模型的可复现框架,采用特定模型与训练策略,可在单台8-GPU服务器数天内完成训练,旨在降低训练门槛以推动相关研究。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18601 2026-07-14 cs.CV 版本更新 96%

Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

Incantation: 自然语言作为多实体视频世界模型的动作接口

Shangwen Zhu, Qianyu Peng, Zhao Pu, Zhilei Shu, Xiangrui Ke, Zhaohu Xing, Zizhao Tong, Zeqing Wang, Xinyu Cui, Zian Zheng, Huangji Wang, Jian Zhao, Yeying Jin, Fan Cheng, Ruili Feng

机构 * SJTU(上海交通大学) NVIDIA Research(英伟达研究) USTC(中国科学技术大学) UCAS(乌兹别克斯坦科学院) NUS(新加坡国立大学) UWaterloo(滑铁卢大学) HKUST(香港理工大学) HKU(香港大学) ZGCA(浙江大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 本研究提出了一种基于自然语言的动作接口,用于多实体视频世界模型,解决了传统接口在细粒度多实体控制和跨实体、跨世界泛化能力上的不足,通过引入自然语言条件化实现了更强大的表达能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31734 2026-07-01 cs.CV 新提交 96%

MemLearner: Learning to Query Context memory for Video World Models

MemLearner: 学习查询上下文记忆用于视频世界模型

Jiwen Yu, Jianxiong Gao, Jianhong Bai, Yiran Qin, Kaiyi Huang, Quande Liu, Xintao Wang, Pengfei Wan, Kun Gai, Xihui Liu

机构 * The University of Hong Kong(香港大学) Fudan University(复旦大学) Zhejiang University(浙江大学) Kuaishou Technology(快手科技)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出基于学习的自适应上下文查询方法MemLearner,利用视频生成模型自身进行上下文记忆查询,解决视频世界模型在遮挡和动态场景下的场景一致性问题。

Comments ECCV 2026, Project Page: https://yujiwen.github.io/memlearner/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21172 2026-06-23 cs.CV 新提交 96%

BadDreamer: Transferable Backdoor Attacks against Video World Models for Autonomous Driving

BadDreamer: 针对自动驾驶视频世界模型的可迁移后门攻击

Zhe Shuai, Xiaopeng Xie, Yikun Zeng

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出BadDreamer,一种针对自动驾驶视频世界模型的可迁移时空后门攻击,通过污染未来帧中的触发-擦除序列,使模型在物理触发出现时幻觉化障碍物消失,进而误导下游动作预测。

Comments 19 pages, 8 figures, 3 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00793 2026-06-09 cs.CV 版本更新 96%

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

MBench: 视频世界模型记忆能力的综合基准

Shengjun Zhang, Zhang Zhang, Simin Huang, Zhenyu Tang, Hanyang Wang, Chensheng Dai, Min Chen, Yifan Li, Yuxin Li, Yingjie Chen, Hao Liu, Chen Li, Jing Lyu, Yueqi Duan

机构 * Tsinghua University(清华大学) WeChat Vision, Tecent Inc.(微信视觉,腾讯公司) Peking University(北京大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出MBench基准,通过实体一致性、环境一致性和因果一致性三个核心维度及其12个子维度,系统评估视频世界模型的长期记忆能力,并揭示现有方法在长期状态保持上的关键局限。

Comments Project Page: https://peanutup.github.io/MBench-project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25077 2026-05-26 cs.CV 96%

WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models

WorldCraft: 从相机导航到交互式视频世界模型中的物体操控

Bohai Gu, Taiyi Wu, Yueyang Yuan, Jian Liu, Xiaocheng Lu, Dazhao Du, Jie Zhang, Jinxiang Lai, Shuai Yang, Xiaotong Zhao, Alan Zhao, Song Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) AI Technology Center, Tencent Video, Tencent(腾讯视频AI技术中心,腾讯) Wuhan University(武汉大学) Peking University(北京大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出WorldCraft框架,通过轨迹控制管道(NWT、SP-LoRA、TASP)将交互式视频世界模型从相机导航扩展到物体级轨迹操控,实现用户指定路径下的物体运动与相机导航共存。

Comments Project page: https://nevsdev.github.io/WorldCraft/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07145 2026-03-25 cs.CV 96%

LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models

LiveWorld: 生成视频世界模型中模拟视线外动态

Zicheng Duan, Jiatong Xia, Zeyu Zhang, Wenbo Zhang, Gengze Zhou, Chenhui Gou, Yefei He, Feng Chen, Xinyu Zhang, Lingqiao Liu

机构 * Adelaide University(阿德莱德大学) The Australian National University(澳大利亚国立大学) Monash University(莫纳什大学) Zhejiang University(浙江大学) University of Auckland(奥克兰大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 LiveWorld提出了一种新框架,解决视频世界模型中视线外动态缺失问题,通过持久化全球状态和监控机制实现持续演化,提升场景一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00310 2026-07-02 cs.CV cs.AI 新提交 95%

RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail

RetailSMV:零售场景中基础视频世界模型的外视角与内视角适应

Amirreza Rouhi, Rajat Aggarwal, Parikshit Sakurikar, Anoop M. Namboodiri, Sashi P. Reddi

机构 * DreamVu

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world models(title);world model(title,abstract)

AI总结 研究零售场景下基础视频世界模型的外视角与内视角适应,发现仅用外视角数据训练的模型在多项指标上优于或等同于联合训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21686 2026-08-06 cs.CV 版本更新 95%

WorldMark: A Unified Benchmark Suite for Interactive Video World Models

WorldMark:交互式视频世界模型的统一基准套件

Xiaojie Xu, Zhengyuan Lin, Kang He, Yukang Feng, Xiaofeng Mao, Yuanyang Yin, Yongtao Ge, Kaipeng Zhang

专题命中 视频世界模型 :world model(title,abstract);world models(title);video world model(title);world model(title,abstract)

AI总结 WorldMark 是交互式视频世界模型的统一基准套件,通过适配器解决动作格式不兼容问题,从多维度评估模型,揭示现有协议未察觉的差异,将发布相关数据与代码。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07687 2026-06-09 cs.CV cs.AI 新提交 94%

What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction

什么使视频世界模型潜在空间与动作相关:预测优于重建

Jewon Yeom, Hanseul Kim, Jeongjae Park, Sungmok Jung, Jaejin Lee, Taesup Kim

机构 * Graduate School of Data Science, Seoul National University(首尔大学数据科学研究生院)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 通过统一探针评估,发现动作相关结构主要由时间视频预训练驱动,而非像素重建保真度,其中视频预训练自监督编码器在视觉保真度和动作预测间取得最佳帕累托权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10104 2026-05-27 cs.CV cs.AI cs.LG 94%

Olaf-World: Orienting Latent Actions for Video World Modeling

Olaf-World: 面向视频世界模型的潜在动作定向

Yuxin Jiang, Yuchao Gu, Ivor W. Tsang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore Research (A STAR), Singapore

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 提出SeqΔ-REPA对齐目标,通过冻结自监督视频编码器的时序特征差异锚定潜在动作,实现无标签视频中可迁移的动作控制世界模型预训练。

Comments ICML 2026. Project page: https://showlab.github.io/Olaf-World/ Code: https://github.com/showlab/Olaf-World

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02473 2026-08-20 cs.CV cs.LG 版本更新 94%

WorldPack: Dynamic Frame Compression for Long-context Video World Modeling

WorldPack:用于长上下文视频世界建模的动态帧压缩

Yuta Oshima, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

机构 * The University of Tokyo(东京大学) Google DeepMind(谷歌DeepMind)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 研究针对长上下文视频世界建模中时空一致生成的挑战,提出WorldPack视频世界模型,通过轨迹打包和几何选择机制,基于3D空间相关性动态分配压缩率,扩展有效上下文,在相关数据集上优于基线,提升空间推理任务表现。

Comments Published in TMLR (09/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02753 2026-06-03 cs.CV cs.AI 94%

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data

MetaWorld: 从单视角视频数据扩展多智能体视频世界模型

Teng Hu, Mingchun Lu, Yating Wang, Jiangning Zhang, Jinkun Hao, Ye Pan, Ran Yi, Lizhuang Ma, Dacheng Tao

机构 * Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 提出MetaWorld框架,通过单目世界状态展开、主体感知世界生成器和世界状态对齐机制,从单视角视频构建多智能体视频世界模型,解决数据稀缺和世界状态对齐问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20083 2026-07-02 cs.CV 新提交 94%

Holo-World: Unified Camera, Object and Weather Control for Video World Model

Holo-World: 视频世界模型的统一相机、物体和天气控制

Xiangchen Yin, Wenzhang Sun, Jiahui Yuan, Zijie Liu, Yinda Chen, Wei Li, Dachun Kai, Chunfeng Wang, Xiaoyan Sun

机构 * University of Science and Technology of China(中国科学技术大学) Li Auto Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究院)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 提出Holo-World,一种从单张图像联合控制相机、物体运动和天气的统一视频世界模型,通过场景适配器和解耦CFG实现世界保持与天气迁移。

Comments Project Page: https://xiangchenyin.github.io/Holo-World Code: https://github.com/XiangchenYin/Holo-World

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09507 2026-06-09 cs.CV 新提交 94%

Prisma-World: Camera-Controllable Multi-Agent Video World Model

Prisma-World: 相机可控的多智能体视频世界模型

Huiqiang Sun, Zhan Peng, Size Wu, Kun Wang, Kang Liao, Dianyi Wang, Xingyu Zeng, Sheng Jin, Yangguang Li, Zhiguo Cao, Ziwei Liu, Wei Li

机构 * School of AIA, HUST(华中科技大学人工智能与自动化学院) S-Lab, NTU(南洋理工大学S-Lab) SenseTime Research(商汤科技研究院) FDU(复旦大学) SUAT(深圳大学) HKU(香港大学) CUHK(香港中文大学)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 提出Prisma-World,通过联合几何感知去噪过程实现多智能体视频生成中的跨视角一致性,支持灵活智能体数量和相机控制。

Comments Project page: https://huiqiang-sun.github.io/prisma-world/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06508 2026-05-26 cs.RO 94%

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy

World-VLA-Loop: 视频世界模型与VLA策略的闭环学习

Xiaokang Liu, Zechen Bai, Hai Ci, Kevin Yuchen Ma, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 提出World-VLA-Loop框架,通过状态感知视频世界模型联合预测未来帧和二元奖励,并采用协同进化范式迭代优化VLA策略,减少对真实环境交互的依赖。

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05138 2026-03-31 cs.CV 94%

VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

VerseCrafter:基于4D几何控制的动态真实视频世界模型

Sixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li, Ying Shan, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) HKU(香港大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 本文提出VerseCrafter,通过4D几何控制生成动态真实视频,相比传统方法更精确控制相机和多物体运动。

Comments Project Page: https://sixiaozheng.github.io/VerseCrafter_page/, Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26037 2026-07-29 cs.CV cs.GR 新提交 94%

Wonder: Video World Model Done Better

Wonder:改进的视频世界模型

Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei

机构 * Adobe Research(Adobe研究院) Johns Hopkins University(约翰·霍普金斯大学)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 Wonder是用于实时相机可控世界探索的通用视频世界模型,通过系统级协同设计,包括新型相机条件设定、高效内存机制等,能合成多样视频,支持视频条件生成,保持长时间的连贯视觉效果。

Comments Project Page: https://wonder-world-model.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13489 2026-08-14 cs.CV cs.RO 新提交 94%

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

DreamX-Phi 1.0:面向机器人操控的动作条件视频世界模型

DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang

机构 * DreamX

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 DreamX-Phi 1.0是面向机器人操控的动作条件视频世界模型,通过几何编码、深度分支等优化,在WorldArena 2.0挑战赛获Track1第一、Track2第二,模型与代码将公开。

Comments Code: https://github.com/AMAP-ML/DreamX-Phi

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15185 2026-05-15 cs.CV cs.AI 94%

Quantitative Video World Model Evaluation for Geometric-Consistency

几何一致性定量视频世界模型评估

Jiaxin Wu, Yihao Pi, Yinling Zhang, Yuheng Li, Xueyan Zou

机构 * Tsinghua University - IEI Lab(清华大学-IEI实验室) UW-Madison(威斯康星大学麦迪逊分校) Adobe Research(Adobe研究院)

专题命中 视频世界模型 :world model(title,abstract);video world model(title);world model(title,abstract);video world model(title)

AI总结 本文提出PDI-Bench框架,通过分割与点跟踪获取物体中心观测,利用单目重建获取3D坐标,计算投影几何残差以评估生成视频的几何一致性,揭示视频生成器的特定失败模式。

Comments 12 pages, 5 figures. Project page : https://pdi-bench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13376 2026-06-18 cs.CV 新提交 93%

MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold

MoVerse: 基于全景高斯支架的实时视频世界建模

Yang Zhou, Ziheng Wang, Yuqin Lu, Haofeng Liu, Jun Liang, Shengfeng He, Jing Li

机构 * South China University of Technology Columbia University Orange Team, Youku Moku-Lab, HUJING Digital Media \& Entertainment Group Singapore Management University

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 提出MoVerse,从单张窄视场图像实时构建可交互漫游的360度全景世界,通过拓扑感知扩散补全视场、全景几何残差预测生成3D高斯支架,并结合双向扩散教师蒸馏为因果自回归学生实现低延迟视频渲染。

Comments Project Page: https://orange-3dv-team.github.io/MoVerse/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20698 2026-06-23 cs.RO 新提交 91%

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model

SafeDojo:基于交互世界模型的安全强化学习用于视觉-语言-动作模型

Kai Tang, Peidong Jia, Zhong Chu, Jixian Wu, Rui Ma, Jiajun Cao, Fangyuan Zhao, Sixiang Chen, Yichen Guo, Xiaowei Chi, Chun-Kai Fan, Kevin Zhang, Jinchang Xu, Fubing Yang, Weishi Mi, Xiaozhu Ju, Jian Tang, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) Nanyang Technological University(南洋理工大学) Hong Kong University of Science and Technology(香港科技大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 视频世界模型 :world model(title,abstract);world model(title,abstract);video world model(abstract);video world model(abstract)

AI总结 提出SafeDojo,首个基于模型的安全强化学习框架,通过交互式视频世界模型进行想象学习安全动作,结合解耦的任务奖励和安全代价信号,在SafeLIBERO和真实机器人上取得最佳安全成功率。

Comments 20 pages, 5 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01060 2026-07-16 cs.RO 版本更新 89%

RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation

RoboWorld: 用于通用机器人策略评估的快速可靠神经模拟器

Byeongguk Jeon, Seonghyeon Ye, JaeHyeok Doo, Sungdong Kim, Minjoon Seo, Hyungmok Son, Kimin Lee

机构 * KAIST(韩国科学技术院) Config

专题命中 视频世界模型 :world model(abstract);world models(abstract);world-model(abstract);video world model(abstract)

AI总结 提出RoboWorld自动化评估流程,结合快速自回归视频世界模型和任务进度感知视觉语言模型评分,通过Step Forcing减少训练-测试不匹配,实现与真实世界评估高度一致。

Comments Project page: https://byeongguks.github.io/RoboWorld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02330 2026-08-03 cs.CV cs.AI cs.LG 版本更新 87%

ActionParty: Multi-Subject Action Binding in Generative Video Games

ActionParty:生成视频游戏中的多主体动作绑定

Alexander Pondaven, Ziyi Wu, Igor Gilitschenski, Philip Torr, Sergey Tulyakov, Fabio Pizzati, Aliaksandr Siarohin

机构 * Snap Research(Snap研究院) University of Oxford(牛津大学) University of Toronto(多伦多大学) MBZUAI(穆罕默德·本·扎耶德人工智能大学)

专题命中 视频世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文提出ActionParty,一种可控制多主体的生成视频游戏世界模型,通过引入主体状态标记和空间偏置机制,提升动作关联准确性与身份一致性。

Comments ECCV 2026 - Project page: https://action-party.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19191 2026-07-22 cs.CV cs.AI cs.LG 新提交 83%

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

ABot-World-0:在单台桌面GPU上进行无限交互式世界展开

Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo

专题命中 视频世界模型 :world model(abstract);video world model(abstract);world model(abstract);video world model(abstract)

AI总结 介绍ABot-World-0这一用于实时长视野闭环交互的动作条件视频世界模型,利用多源数据学习世界动态。通过多种技术提炼模型,设计控制界面与部署堆栈,在单台桌面GPU上实现高效视频流传输,实验验证其有竞争力的可控性和世界演变能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09038 2025-02-28 cs.CV cs.AI cs.GR cs.LG 83%

Do generative video models understand physical principles?

Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, Robert Geirhos

专题命中 视频世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07600 2025-05-22 cs.CV cs.RO 81%

PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning

Angel Villar-Corrales, Sven Behnke

机构 * Autonomous Intelligent Systems, Computer Science Institute VI – Intelligent Systems(智能系统) Robotics, Center for Robotics(机器人学) the Lamarr Institute for Machine Learning(拉马尔机器学习研究所) Artificial Intelligence, University of Bonn, Germany(人工智能,波恩大学,德国)

专题命中 视频世界模型 :latent dynamics(title);world model(abstract);world model(abstract);分类 cs.CV、cs.RO

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14718 2025-07-15 cs.CV 81%

A Survey on Future Frame Synthesis: Bridging Deterministic and Generative Approaches

Ruibo Ming, Zhewei Huang, Jingwei Wu, Zhuoxuan Ju, Daxin Jiang, Jianming Hu, Lihui Peng, Shuchang Zhou

机构 * Tsinghua University(清华大学) StepFun Peking University(北京大学) Megvii Technology(旷视科技)

专题命中 视频世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments TMLR 2025/07

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14948 2025-05-22 cs.CV cs.AI cs.LG 73%

Programmatic Video Prediction Using Large Language Models

Hao Tang, Kevin Ellis, Suhas Lohit, Michael J. Jones, Moitreya Chatterjee

机构 * Mitsubishi Electric Research Laboratories (MERL)(三菱电机研究实验室(MERL))

专题命中 视频世界模型 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏