arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 4334 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4334 篇

2605.26379 2026-05-27 stat.ML cs.LG 90%

When Does LeJEPA Learn a World Model?

LeJEPA 何时学习世界模型?

David Klindt, Yann LeCun, Randall Balestriero

机构 * Cold Spring Harbor Laboratory(冷泉港实验室) New York University(纽约大学) Brown University(布朗大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文证明 LeJEPA(对齐加高斯正则化)在潜变量服从平稳加性噪声演化的世界中能够线性恢复潜变量(线性可识别性),并指出高斯分布是唯一保证该性质的潜分布,同时验证了近似可识别性和最优规划能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24375 2026-05-26 cs.AI 90%

Distilling Game Code World Model Generation into Lightweight Large Language Models

将游戏代码世界模型生成蒸馏到轻量级大型语言模型

Tyrone Serapio, Arjun Prakash, Haoyang Xu, Kevin Wang, Amy Greenwald

机构 * Brown University(布朗大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 研究通过后训练将游戏代码世界模型生成能力蒸馏到小型模型,采用监督微调和带可验证奖励的强化学习提升生成代码的语法正确性和规则遵循性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23025 2026-05-25 cs.LG 90%

World Machine: Towards Generative World Modeling for Time-Series

世界机器:面向时间序列的生成式世界建模

Elton Cardoso do Nascimento, Alexandre da Silva Simões, Esther Luna Colombini, Ricardo Ribeiro Gudwin, Paula Dornhofer Paro Costa

机构 * Universidade Estadual de Campinas (UNICAMP)(坎皮纳斯州立大学) Universidade Estadual Paulista (UNESP)(保罗斯州立大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 提出基于Transformer的潜在状态架构World Machine,用于时间序列生成式世界建模,在合成数据集Toy1D上验证了其超越传统Transformer的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11836 2026-05-22 cs.AI 90%

Finite Automata Extraction: Low-data World Model Learning as Programs from Gameplay Video

有限自动机提取:从游戏录像中学习低数据世界模型作为程序

Dave Goel, Matthew Guzdial, Anurag Sarkar

机构 * Department of Computing Science, Alberta Machine Intelligence Institute (Amii), University of Alberta(计算科学系,阿尔伯塔机器智能研究所(Amii),阿尔伯塔大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出了一种名为有限自动机提取(FAE)的方法,通过一种新的领域特定语言(DSL)Retro Coder,从游戏录像中学习神经符号世界模型,相较于以往的方法,FAE能够更精确地建模环境并生成更通用的代码。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02193 2026-05-22 cs.AI 90%

From monoliths to modules: Decomposing transducers for efficient world modelling

从整体到模块:分解转换器以实现高效的world建模

Alexander Boyd, Franz Nowak, David Hyland, Manuel Baltieri, Fernando E. Rosas

机构 * Department of Informatics, University of Sussex(Sussex大学信息学院) Beyond Institute for Theoretical Science (BITS)(理论科学研究所) ETH Zürich(苏黎世联邦理工学院) Principles of Intelligent Behaviour in Biological and Social Systems (PIBBSS)(生物和社会系统智能行为原理研究所) Department of Computer Science, University of Oxford(牛津大学计算机科学系) Araya Inc.(Araya公司) Sussex AI and Sussex Centre for Consciousness Science, University of Sussex(Sussex大学人工智能与意识科学中心) Centre for Complexity Science and Center for Psychedelic Research, Department of Brain Sciences, Imperial College London(复杂科学中心和迷幻研究中心,伦敦帝国理工学院脑科学系) Center for Eudaimonia and Human Flourishing, University of Oxford(幸福与人类繁荣中心,牛津大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出了一种分解复杂world建模的方法,通过转换器框架将世界模型分解为多个模块,从而提高计算效率并支持分布式推理,为AI安全和现实应用提供基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16530 2026-05-21 cs.CV 90%

SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation

SWoMo:用于白内障手术模拟的神经符号世界模型

Ssharvien Kumar Sivakumar, Akwele Johnson, Anirudh Dhingra, Yannik Frisch, Ghazal Ghazaei, Anirban Mukhopadhyay

机构 * Technical University Darmstadt(德累斯顿技术大学) Carl Zeiss AG(蔡司股份有限公司) AICM, Medical Faculty of Heidelberg University(海德堡大学医学院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出SWoMo,一种用于白内障手术模拟的神经符号世界模型,通过分离运动生成与视觉真实性,结合规则基模拟器和场景图表示来建模运动动态和工具-组织交互,同时使用扩散模型生成逼真的视觉效果,从而提升手术模拟的真实性和临床适用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17451 2026-05-19 cs.CV 90%

DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking

DeTrack:一种无人机具身跟踪的基准及海拔感知双世界模型

Guyue Hu, Haoming Liu, Siyuan Song, Chenglong Li, Feng Chen, Jin Tang

机构 * Hefei Si Valley Technology Development Co., Ltd(合肥蜀山科技发展有限公司) Institute of Embodied Intelligence, Anhui University(embodied intelligence研究院,安徽大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出DeTrack任务,要求无人机在交互式3D环境中利用在线自体观察和主动飞行控制进行目标跟踪,并提出AaDWorlds框架以解决海拔相关的可见性与飞行安全矛盾。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08398 2026-05-18 cs.CV 90%

VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?

VideoVerse: 你的T2V生成器有世界模型能力来合成视频吗?

Zeqing Wang, Xinyu Wei, Bairui Li, Zhen Guo, Jinrui Zhang, Hongyang Wei, Keze Wang, Lei Zhang

机构 * Sun Yat-sen University(中山大学) Hong Kong Polytechnic University(香港理工大学) Tsinghua University(清华大学) OPPO Research Institute(OPPO研究院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 VideoVerse通过评估T2V模型对复杂时间因果关系和世界知识的理解能力,揭示现有模型与理想世界建模能力的差距。

Comments 26 Pages, 10 Figures, 14 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15178 2026-05-15 cs.CV 90%

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

SANA-WM:高效分钟级世界建模的混合线性扩散变换器

Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, Junsong Chen, Jincheng Yu, Tong He, Song Han, Enze Xie

机构 * NVIDIA(英伟达)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 SANA-WM通过混合线性注意力、双分支相机控制等设计,实现高效分钟级视频生成,提升效率与视觉质量。

Comments https://nvlabs.github.io/Sana/WM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07326 2026-05-11 cs.CV 90%

GEM: Generating LiDAR World Model via Deformable Mamba

GEM: 通过可变形Mamba生成LiDAR世界模型

Yang Wu, Zhaojiang Liu, Qiang Meng, Youquan Liu, Renliang Weng, Jianjun Qian, Jian Yang, Jin Xie

机构 * NJU(南京大学) SJTU(上海交通大学) FDU(福建师范大学) NTU(南京理工大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 GEM通过可变形Mamba架构提升LiDAR世界模型的逼真度与想象力,解决点云无序和动态静态区分难题,实现空间时间感知与自主探索能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06337 2026-05-08 cs.CV 90%

Earth-o1: A Grid-free Observation-native Atmospheric World Model

Earth-o1:一种无网格的观测本征大气世界模型

Junchao Gong, Kaiyi Xu, Wangxu Wei, Siwei Tu, Jingyi Xu, Zili Liu, Hang Fan, Zhiwang Zhou, Tao Han, Yi Xiao, Xinyu Gu, Zhangrui Li, Wenlong Zhang, Hao Chen, Xiaokang Yang, Yaqiang Wang, Lijing Cheng, Pierre Gentine, Wanli Ouyang, Feng Zhang, Zhe-Min Tan, Bowen Zhou, Fenghua Ling, Ben Fei, Lei Bai

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Department of Information Engineering(信息工程系) School of Electronic Information and Electrical Engineering(电子信息与电气工程学院) Department of Atmospheric and Oceanic Sciences(大气与海洋科学系) School of Information Science and Technology(信息科学与技术学院) State Key Laboratory of Earth System Numerical Modeling and Application, Institute of Atmospheric Physics(地球系统数值模拟与应用国家重点实验室,大气物理研究所) College of Computer Science and Artificial Intelligence(计算机科学与人工智能学院) Department of Earth and Environmental Engineering(地球与环境工程系) Chinese Academy of Meteorological Sciences(中国气象科学研究院) School of Atmospheric Sciences(大气科学学院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 Earth-o1通过直接学习无网格观测数据中的连续三维物理演变,实现对大气状态的自主时空推进,无需传统数值求解器,其预测精度可与现有物理框架相媲美。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26848 2026-05-04 cs.RO 90%

STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation

STARRY:面向机器人操作的时空动作中心世界建模

Yuxuan Tian, Yurun Jin, Bin Yu, Yukun Shi, Hao Wu, Chi Harold Liu, Kai Chen, Cong Huang

机构 * Beijing Institute of Technology(北京理工大学) Zhongguancun Academy(中关村学院) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院) University of Science and Technology of China(中国科学技术大学) Harbin Institute of Technology(哈尔滨工业大学) East China Normal University(华东师范大学) DeepCybo

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 STARRY通过统一扩散过程联合去噪未来时空潜在表示和动作,提升机器人操作中时空协调的精度与成功率。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16824 2026-04-21 cs.CR cs.AI 90%

SafeDream: Safety World Model for Proactive Early Jailbreak Detection

SafeDream: 用于主动早期对抗检测的安全世界模型

Bo Yan, Weikai Lin, Yada Zhu, Song Wang

机构 * University of Central Florida(中央佛罗里达大学) University of Rochester(罗切斯特大学) IBM Research(IBM研究院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 本文提出SafeDream,一种轻量级世界模型框架,通过安全状态模型、CUSUM检测和对比想象组件,实现早期对抗检测,提升检测及时性并降低误报率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14268 2026-04-17 cs.CV 90%

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

HY-World 2.0:一个多模态世界模型用于重建、生成和模拟3D世界

Team HY-World, Chenjie Cao, Xuhui Zuo, Zhenwei Wang, Yisu Zhang, Junta Wu, Zhenyang Liu, Yuning Gong, Yang Liu, Bo Yuan, Chao Zhang, Coopers Li, Dongyuan Guo, Fan Yang, Haiyu Zhang, Hang Cao, Jianchen Zhu, Jiaxin Lin, Jie Xiao, Jihong Zhang, Junlin Yu, Lei Wang, Lifu Wang, Lilin Wang, Linus, Minghui Chen, Peng He, Penghao Zhao, Qi Chen, Rui Chen, Rui Shao, Sicong Liu, Wangchen Qin, Xiaochuan Niu, Xiang Yuan, Yi Sun, Yifei Tang, Yifu Sun, Yihang Lian, Yonghao Tan, Yuhong Liu, Yuyang Yin, Zhiyuan Min, Tengfei Wang, Chunchao Guo

机构 * Tencent(腾讯)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 HY-World 2.0通过多模态输入生成高保真3D场景,改进了先前版本,引入了多项创新以提升全景真实性、3D场景理解和规划能力,并升级了WorldStereo和WorldMirror模型,实现了3D世界交互探索。

Comments Project Page: https://3d-models.hunyuan.tencent.com/world/ ; Code: https://github.com/Tencent-Hunyuan/HY-World-2.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08995 2026-04-14 cs.CV 90%

Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory

矩阵游戏3.0:具有长时程记忆的实时和流式交互世界模型

Zile Wang, Zexiang Liu, Jiaxing Li, Kaichen Huang, Baixin Xu, Fei Kang, Mengyin An, Peiyu Wang, Biao Jiang, Yichen Wei, Yidan Xietian, Jiangbo Pei, Liang Hu, Boyi Jiang, Hua Xue, Zidong Wang, Haofeng Sun, Wei Li, Wanli Ouyang, Xianglong He, Yang Liu, Yangguang Li, Yahui Zhou

机构 * Skywork AI(天工AI)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 矩阵游戏3.0通过改进数据、模型和推理方法,实现了720p实时长视频生成,支持长时程时空一致性,提升了生成效率和质量。

Comments Project page: https://matrix-game-v3.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13009 2026-04-08 cs.CV 90%

Matrix-game 2.0: An open-source real-time and streaming interactive world model

矩阵游戏2.0:一个开源的实时和流式交互世界模型

Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, Baixin Xu, Hao-Xiang Guo, Kaixiong Gong, Size Wu, Wei Li, Xuchen Song, Yang Liu, Yangguang Li, Yahui Zhou

机构 * Skywork AI(天工AI)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出Matrix-Game 2.0,通过少步自回归扩散生成长视频,解决传统交互世界模型实时性差的问题,实现25FPS高速生成。

Comments Project Page: https://matrix-game-v2.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12655 2026-03-16 cs.CV 90%

VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model

VGGT-World:将VGGT转变为一个自回归的几何世界模型

Xiangyu Sun, Shijie Wang, Fengyi Zhang, Lin Liu, Caiyan Jia, Ziying Song, Zi Huang, Yadan Luo

机构 * UQMM Lab, The University of Queensland(昆士兰大学UQMM实验室) Beijing Jiaotong University(北京交通大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出VGGT-World,通过自回归方法预测冻结几何基础模型特征的时空演变,提升深度预测性能并提高效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07562 2026-03-10 cs.CV 90%

Brain-WM: Brain Glioblastoma World Model

Brain-WM: 脑部胶质瘤世界模型

Chenhui Wang, Boyun Zheng, Liuxin Bao, Zhihao Peng, Peter Y. M. Woo, Hongming Shan, Yixuan Yuan

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 Brain-WM通过统一治疗预测与MRI生成,实现了肿瘤与治疗的共进化动态建模,提升了治疗计划的准确率和MRI生成的质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11291 2026-03-05 cs.RO 90%

H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model

H-WM:由分层世界模型引导的机器人任务和运动规划

Jinbang Huang, Wenyuan Chen, Zhiyuan Li, Oscar Pang, Xiao Hu, Lingfeng Zhang, Yuanzhao Hu, Zhanguang Zhang, Mark Coates, Tongtong Cao, Xingyue Quan, Yingxue Zhang

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) University of Toronto(多伦多大学) University of British Columbia(不列颠哥伦比亚大学) McGill University(麦吉尔大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 H-WM通过结合逻辑和视觉世界模型,实现机器人任务和运动规划中的鲁棒长视界推理与稳定引导。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12063 2026-02-17 cs.RO 90%

VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model

VLAW: 视觉-语言-动作策略与世界模型的迭代共改进

Yanjiang Guo, Tony Lee, Lucy Xiaoyang Shi, Jianyu Chen, Percy Liang, Chelsea Finn

机构 * Stanford University(斯坦福大学) Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出通过迭代改进视觉-语言-动作策略与世界模型,提升真实机器人任务的性能和可靠性。

Comments Project Page: https://sites.google.com/view/vlaw-arxiv

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10884 2026-02-12 cs.CV 90%

ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving

ResWorld: 用于端到端自动驾驶的时序残差世界模型

Jinqing Zhang, Zehua Fu, Zelin Xu, Wenying Dai, Qingjie Liu, Yunhong Wang

机构 * State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,北京航空航天大学) Zhongguancun Laboratory, Beijing, China(中关村实验室) Beijing Jingwei Hirain Technologies Co., Inc.(北京京wei Hirain科技有限公司) Hangzhou Innovation Institute, Beihang University, Hangzhou, China(杭州创新研究院,北京航空航天大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 ResWorld通过时序残差世界模型和未来引导轨迹细化模块,提升端到端自动驾驶的规划性能。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23429 2026-02-11 cs.CV 90%

Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model

Hunyuan-GameCraft-2:基于指令的交互式游戏世界模型

Junshu Tang, Jiacheng Liu, Jiaqi Li, Longhuang Wu, Haoyu Yang, Penghao Zhao, Siruis Gong, Xiang Yuan, Shuai Shao, Linfeng Zhang, Qinglin Lu

机构 * Tencent Hunyuan(腾讯 Hunyuan)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 Hunyuan-GameCraft-2通过自然语言提示等多模态交互方式,实现更灵活的生成游戏世界建模,提升交互性和因果一致性。

Comments Technical Report, Project page:https://hunyuan-gamecraft-2.github.io/, Demo:https://hunyuan.tencent.com/game/game-craft

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02002 2026-02-03 cs.CV 90%

UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving

UniDriveDreamer: 一种用于自动驾驶的单阶段多模态世界模型

Guosheng Zhao, Yaozeng Wang, Xiaofeng Wang, Zheng Zhu, Tingdong Yu, Guan Huang, Yongchen Zai, Ji Jiao, Changliang Xue, Xiaole Wang, Zhen Yang, Futang Zhu, Xingang Wang

机构 * GigaAI CASIA BYD

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 UniDriveDreamer是一种用于自动驾驶的单阶段多模态世界模型,通过统一的多模态生成方法提升视频和激光雷达数据合成的性能。

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18620 2026-01-27 cs.LG 90%

CASSANDRA: Programmatic and Probabilistic Learning and Inference for Stochastic World Modeling

CASSANDRA:面向随机世界建模的程序化与概率学习与推理

Panagiotis Lymperopoulos, Abhiramon Rajasekharan, Ian Berlot-Attwell, Stéphane Aroca-Ouellette, Kaheer Suleman

机构 * Skyfall AI Tufts University(塔夫茨大学) University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Toronto(多伦多大学) Vector Instiute(Vector研究所)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 CASSANDRA通过结合LLM的知识先验和概率图模型结构学习,提升随机世界建模中的转移预测与规划能力。

Comments 28 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12277 2026-01-21 cs.RO 90%

An Efficient and Multi-Modal Navigation System with One-Step World Model

一种高效且多模态的导航系统与一步世界模型

Wangtian Shen, Ziyang Meng, Jinming Ma, Mingliang Zhou, Diyun Xiang

机构 * Tsinghua University(清华大学) Xiaomi (China)(小米(中国))

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出了一种高效多模态导航系统,通过一步世界模型和3D U-Net骨干网络,提升导航效率和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04035 2026-01-08 cs.AI 90%

MobileDreamer: Generative Sketch World Model for GUI Agent

MobileDreamer: 用于GUI代理的生成式草图世界模型

Yilin Cao, Yufeng Zhong, Zhixiong Zeng, Liming Zheng, Jing Huang, Haibo Qiu, Peng Shi, Wenji Mao, Wan Guanglu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Meituan(美团)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 MobileDreamer通过生成式草图世界模型和rollout想象策略,提升GUI代理在长周期任务中的决策能力,任务成功率提升5.25%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00051 2026-01-05 cs.CV 90%

TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model

TeleWorld:面向动态多模态合成的4D世界模型

Yabo Chen, Yuanzhi Liang, Jiepeng Wang, Tingxi Chen, Junfei Cheng, Zixiao Gu, Yuyang Huang, Zicheng Jiang, Wei Li, Tian Li, Weichen Li, Zuoxin Li, Guangce Liu, Jialun Liu, Junqi Liu, Haoyuan Wang, Qizhen Weng, Xuan'er Wu, Xunzhi Xiang, Xiaoyan Yang, Xin Zhang, Shiwen Zhang, Junyu Zhou, Chengcheng Zhou, Haibin Huang, Chi Zhang, Xuelong Li

机构 * TeleWorld Team(TeleWorld团队)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 TeleWorld提出了一种实时多模态4D世界建模框架,通过生成-重建-引导范式实现动态场景重建与长期记忆,提升世界模型的交互性和计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20425 2026-01-01 cs.RO 90%

OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation

OSVI-WM:通过世界模型引导轨迹生成实现单次视觉模仿

Raktim Gautam Goswami, Prashanth Krishnamurthy, Yann LeCun, Farshad Khorrami

机构 * New York University Tandon School of Engineering(纽约大学Tandon工程学院) New York University Courant Institute of Mathematical Sciences(纽约大学Courant数学科学研究所) Meta-FAIR

专题命中 通用世界模型 :world-model(title,abstract);world-model(title,abstract);world model(abstract);world model(abstract)

AI总结 OSVI-WM通过世界模型引导轨迹生成,实现对未见任务的单次视觉模仿学习,提升机器人执行复杂任务的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08139 2025-12-23 cs.IT cs.LG math.IT 90%

SCA-LLM: Spectral-Attentive LLM-Based Wireless World Modeling for Agentic Communications

SCA-LLM:基于频谱-注意力的LLM无线世界建模用于智能通信

Ke He, Le He, Lisheng Fan, Xianfu Lei, Thang X. Vu, George K. Karagiannidis, Symeon Chatzinotas

机构 * Interdisciplinary Centre for Security, Reliability and Trust (SnT), University of Luxembourg(安全、可靠性与信任跨学科研究中心(SnT),卢森堡大学) School of Computer Science of Guangzhou University(广州大学计算机科学学院) School of Information Science and Technology, Institute of Mobile Communications, Southwest Jiaotong University(信息科学与技术学院,移动通信研究所,西南交通大学) Department of Electrical and Computer Engineering, Aristotle University of Thessaloniki(电气与计算机工程系,塞萨洛尼基阿瑞斯托大学) Cyber Security Systems and Applied AI Research Center, Lebanese American University (LAU)(网络安全与应用人工智能研究中心,黎巴嫩美国大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 SCA-LLM通过频谱-注意力适配器将信道状态信息与LLM结合,实现无线世界建模,提升预测性能和零样本泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17152 2025-12-22 cs.CV 90%

PhysFire-WM: A Physics-Informed World Model for Emulating Fire Spread Dynamics

PhysFire-WM: 一种融合物理信息的世界模型用于模拟火灾扩散动力学

Nan Zhou, Huandong Wang, Jiahao Li, Yang Li, Xiao-Ping Zhang, Yong Li, Xinlei Chen

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 PhysFire-WM通过融合物理信息和跨任务协作训练策略,提升火灾扩散预测的物理真实性和几何准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏