arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-02-11 至 2026-02-11 共收录 8 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 6 篇

2602.03213 2026-02-11 cs.CV 96%

ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask

ConsisDrive: 用于视频生成的实例掩码驾驶世界模型

Zhuoran Yang, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);driving world model(title,abstract);world model(title,abstract)

AI总结 ConsisDrive通过实例掩码注意力和损失机制提升驾驶视频生成质量及自动驾驶性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10044 2026-02-11 cs.LG cs.AI cs.SY eess.SY 94%

Optimistic World Models: Efficient Exploration in Model-Based Deep Reinforcement Learning

乐观世界模型:基于模型的深度强化学习中的高效探索

Akshay Mete, Shahid Aamir Sheikh, Tzu-Hsiang Lin, Dileep Kalathil, P. R. Kumar

机构 * Department of Electrical Computer Engineering, Texas A\&M University, College Station, Texas, USA

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出乐观世界模型,通过引入乐观动态损失提升深度强化学习中的高效探索能力,显著提高样本效率和累积回报。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23429 2026-02-11 cs.CV 90%

Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model

Hunyuan-GameCraft-2:基于指令的交互式游戏世界模型

Junshu Tang, Jiacheng Liu, Jiaqi Li, Longhuang Wu, Haoyu Yang, Penghao Zhao, Siruis Gong, Xiang Yuan, Shuai Shao, Linfeng Zhang, Qinglin Lu

机构 * Tencent Hunyuan(腾讯 Hunyuan)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 Hunyuan-GameCraft-2通过自然语言提示等多模态交互方式,实现更灵活的生成游戏世界建模,提升交互性和因果一致性。

Comments Technical Report, Project page:https://hunyuan-gamecraft-2.github.io/, Demo:https://hunyuan.tencent.com/game/game-craft

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09856 2026-02-11 cs.CV cs.AI cs.CL cs.HC 88%

Code2World: A GUI World Model via Renderable Code Generation

Code2World: 通过可渲染代码生成实现一个GUI世界模型

Yuhao Zheng, Li'an Zhong, Yi Wang, Rui Dai, Kaikui Liu, Xiangxiang Chu, Linyuan Lv, Philip Torr, Kevin Qinghong Lin

机构 * University of Science(科学大学) AMAP, Alibaba Group(AMAP,阿里巴巴集团) University of Oxford(牛津大学) Sun Yat-sen University(中山大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI、cs.CV

AI总结 Code2World通过可渲染代码生成实现高保真UI预测,提升Android导航性能

Comments github: https://github.com/AMAP-ML/Code2World project page: https://amap-ml.github.io/Code2World/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21872 2026-02-11 eess.IV cs.LG 69%

Targeted Unlearning Using Perturbed Sign Gradient Methods With Applications On Medical Images

利用扰动符号梯度方法的目标遗忘与医疗图像应用

George R. Nahass, Zhu Wang, Homa Rashidisabet, Won Hwa Kim, Sasha Hubschman, Jeffrey C. Peterson, Chad A. Purnell, Pete Setabutr, Ann Q. Tran, Darvin Yi, Sathya N. Ravi

机构 * Department of Biomedical Engineering(生物医学工程系) Department of Ophthalmology(眼科学系) University of Illinois Chicago(伊利诺伊大学芝加哥分校) Department of Computer Science(计算机科学系) Computer Science and Engineering(计算机科学与工程) Pohang University of Science and Technology, South Korea(韩国釜山科学技术大学) Department of Plastic and Reconstructive Surgery(整形外科与重建外科系)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 本文提出了一种基于扰动符号梯度的方法,用于医疗图像中的目标遗忘,通过可调损失设计和模型组合策略,在遗忘与保留之间取得平衡,优于现有基线方法。

Comments 39 pages, 12 figures, 11 tables, 3 algorithms

Journal ref Transactions on Machine Learning Research 2025, https://openreview.net/forum?id=XE0bJg6sQN

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24654 2026-02-11 cs.CL 67%

Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment

在虚拟临床环境中进化交互诊断代理

Pengcheng Qiu, Chaoyi Wu, Junwei Liu, Qiaoyu Zheng, Yusheng Liao, Haowen Wang, Yun Yue, Qianrui Fan, Shuai Zhen, Jian Wang, Jinjie Gu, Yanfeng Wang, Ya Zhang, Weidi Xie

专题命中 通用世界模型 :world model(abstract);world model(abstract)

AI总结 本研究提出DiagAgent,通过强化学习在虚拟临床环境中训练,实现多轮交互式诊断,显著提升诊断准确性和检查推荐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身与机器人 2 篇

2509.21797 2026-02-11 cs.CV 86%

MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation

MoWM: 通过潜在到像素特征调制的混合世界模型实现具身规划

Yangcheng Yu, Xin Jin, Yu Shang, Xin Zhang, Haisheng Su, Wei Wu, Yong Li

机构 * Tsinghua University(清华大学) Manifold AI Shanghai Jiao Tong University(上海交通大学)

专题命中 具身与机器人 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 MoWM通过融合潜在世界模型与像素特征,提升具身规划中动作解码的精度与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10102 2026-02-11 cs.CV 56%

VideoWorld 2: Learning Transferable Knowledge from Real-world Videos

VideoWorld 2: 从真实世界视频中学习可迁移的知识

Zhongwei Ren, Yunchao Wei, Xiao Yu, Guixun Luo, Yao Zhao, Bingyi Kang, Jiashi Feng, Xiaojie Jin

机构 * Beijing Jiaotong University(北京交通大学)

专题命中 具身与机器人 :latent dynamics(abstract);分类 cs.CV;dynamics model(abstract)

AI总结 VideoWorld 2通过动态增强的潜在动态模型从真实世界视频中学习可迁移知识,显著提升了任务成功率和长周期推理能力。

Comments Code and models are released at: https://maverickren.github.io/VideoWorld2.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏