arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6497 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 具身与机器人 510 篇

2005.00227 2020-05-04 cs.RO cs.LG 50%

Learning Compliance Adaptation in Contact-Rich Manipulation

Jianfeng Gao, You Zhou, Tamim Asfour

专题命中 具身与机器人 :分类 cs.LG、cs.RO;predictive model(abstract);predictive models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.13194 2020-04-29 cs.RO 50%

Learning for Microrobot Exploration: Model-based Locomotion, Sparse-robust Navigation, and Low-power Deep Classification

Nathan O. Lambert, Farhan Toddywala, Brian Liao, Eric Zhu, Lydia Lee, Kristofer S. J. Pister

专题命中 具身与机器人 :model-based reinforcement learning(abstract);分类 cs.RO

Comments 6 pages; 2 pages appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.06489 2020-03-10 cs.RO 50%

Multi-Fidelity Reinforcement Learning with Gaussian Processes

Varun Suryan, Nahush Gondhalekar, Pratap Tokekar

专题命中 具身与机器人 :model-based RL(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.08116 2019-12-18 cs.RO cs.NE 50%

When Your Robot Breaks: Active Learning During Plant Failure

Mariah Schrum, Matthew Gombolay

专题命中 具身与机器人 :latent dynamics(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.06833 2019-11-19 cs.LG cs.AI cs.RO stat.ML 50%

Improved Exploration through Latent Trajectory Optimization in Deep Deterministic Policy Gradient

Kevin Sebastian Luck, Mel Vecerik, Simon Stepputtis, Heni Ben Amor, Jonathan Scholz

专题命中 具身与机器人 :分类 cs.AI、cs.LG、cs.RO;dynamics model(abstract)

Comments Accepted for IROS 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.10251 2019-11-04 cs.LG stat.ML 50%

Multi-Agent Reinforcement Learning with Multi-Step Generative Models

Orr Krupnik, Igor Mordatch, Aviv Tamar

专题命中 具身与机器人 :model-based reinforcement learning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.03710 2019-06-11 cs.LG cs.AI cs.RO stat.ML 50%

Curiosity-Driven Multi-Criteria Hindsight Experience Replay

John B. Lanier, Stephen McAleer, Pierre Baldi

专题命中 具身与机器人 :分类 cs.AI、cs.LG、cs.RO;dynamics model(abstract)

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.04411 2019-05-14 cs.RO cs.CV cs.LG 50%

Learning Robotic Manipulation through Visual Planning and Acting

Angelina Wang, Thanard Kurutach, Kara Liu, Pieter Abbeel, Aviv Tamar

专题命中 具身与机器人 :分类 cs.LG、cs.CV、cs.RO;dynamics model(abstract)

Comments RSS 2019. Website at https://sites.google.com/berkeley.edu/vpa/home

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.09318 2019-01-08 cs.LG stat.ML 50%

Floyd-Warshall Reinforcement Learning: Learning from Past Experiences to Reach New Goals

Vikas Dhiman, Shurjo Banerjee, Jeffrey M. Siskind, Jason J. Corso

专题命中 具身与机器人 :model-based RL(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.11074 2018-11-26 cs.AI 50%

Robot Representation and Reasoning with Knowledge from Reinforcement Learning

Keting Lu, Shiqi Zhang, Peter Stone, Xiaoping Chen

专题命中 具身与机器人 :model-based RL(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.07167 2018-10-17 cs.RO cs.AI cs.LG 50%

Composable Action-Conditioned Predictors: Flexible Off-Policy Learning for Robot Navigation

Gregory Kahn, Adam Villaflor, Pieter Abbeel, Sergey Levine

专题命中 具身与机器人 :分类 cs.AI、cs.LG、cs.RO;predictive model(abstract)

Comments Accepted to the Conference on Robot Learning (CoRL) 2018. Video at https://youtu.be/lOLT7zifEkg

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.07551 2018-07-10 stat.ML cs.LG 50%

Meta Reinforcement Learning with Latent Variable Gaussian Processes

Steindór Sæmundsson, Katja Hofmann, Marc Peter Deisenroth

专题命中 具身与机器人 :model-based reinforcement learning(abstract);分类 cs.LG

Comments 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.00109 2018-06-04 cs.RO cs.LG 50%

Probabilistically Safe Robot Planning with Confidence-Based Human Predictions

Jaime F. Fisac, Andrea Bajcsy, Sylvia L. Herbert, David Fridovich-Keil, Steven Wang, Claire J. Tomlin, Anca D. Dragan

专题命中 具身与机器人 :分类 cs.LG、cs.RO;predictive model(abstract);predictive models(abstract)

Comments Robotics Science and Systems (RSS) 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.00196 2018-03-02 cs.RO cs.AI cs.LG stat.ML 50%

Learning Flexible and Reusable Locomotion Primitives for a Microrobot

Brian Yang, Grant Wang, Roberto Calandra, Daniel Contreras, Sergey Levine, Kristofer Pister

专题命中 具身与机器人 :分类 cs.AI、cs.LG、cs.RO;dynamics model(abstract)

Comments 8 pages. Accepted at RAL+ICRA2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.08560 2017-06-28 cs.RO 50%

Skill Learning by Autonomous Robotic Playing using Active Learning and Creativity

Simon Hangl, Vedran Dunjko, Hans J. Briegel, Justus Piater

专题命中 具身与机器人 :environment model(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.02018 2017-03-07 cs.CV cs.LG cs.RO 50%

Combining Self-Supervised Learning and Imitation for Vision-Based Rope Manipulation

Ashvin Nair, Dian Chen, Pulkit Agrawal, Phillip Isola, Pieter Abbeel, Jitendra Malik, Sergey Levine

专题命中 具身与机器人 :分类 cs.LG、cs.CV、cs.RO;dynamics model(abstract)

Comments 8 pages, accepted to International Conference on Robotics and Automation (ICRA) 2017

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 自动驾驶 144 篇

2606.06014 2026-06-05 cs.AI cs.RO 95%

PLAN-S: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models

PLAN-S:通过潜在风格动态桥接规划以实现自动驾驶世界模型

Xiaoyun Qiu, Jingtao He, Yijie Chen, Yusong Huang, Haotian Wang, Yixuan Wang, Xinhu Zheng

机构 * Intelligent Transportation Thrust, Systems Hub, and Center of Seamless Connectivity & Connected Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(智能交通 thrust、系统中心及无缝连接与智能连接研究院,香港科学与技术大学(广州))

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);driving world model(title);world model(title,abstract)

AI总结 提出PLAN-S框架,通过从潜在表示解码风格条件语义成本图,解决自动驾驶中潜在世界模型规划的可控性问题,在nuScenes和NAVSIM上降低了碰撞率并提升了驾驶性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09059 2026-04-13 cs.CV cs.AI 94%

Learning Vision-Language-Action World Models for Autonomous Driving

学习视觉-语言-动作世界模型以实现自动驾驶

Guoqing Wang, Pin Tang, Xiangxuan Ren, Guodongfang Zhao, Bailan Feng, Chao Ma

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(上海交通大学人工智能研究院教育部人工智能重点实验室) Central Research Institute, Huawei(华为中央研究院)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出VLA-World模型,通过结合预测想象与反思推理提升自动驾驶的前瞻性。该模型利用生成的轨迹引导图像生成,并通过反思优化轨迹预测,实验表明其在规划和未来场景生成任务中优于现有方法。

Comments Accepted by CVPR2026 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14696 2026-05-15 cs.CV 94%

EponaV2: Driving World Model with Comprehensive Future Reasoning

EponaV2:通过全面的未来推理驱动世界模型

Jiawei Xu, Zhizhou Zhong, Zhijian Shu, Mingkai Jia, Mingxiao Li, Jia-Wang Bian, Qian Zhang, Kaicheng Zhang, Jin Xie, Jian Yang, Wei Yin

机构 * PCA Lab, VCIP, College of Computer Science, Nankai University(PCA实验室、VCIP、计算机科学学院、南开大学) Horizon Robotics HKUST(香港科技大学) NJUPT(南京工程大学) NTU(国立台湾大学) Anyverse School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院、南京大学)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 本文提出EponaV2,一种新的驾驶世界模型范式,通过全面的未来推理实现高质量规划。模型通过预测更全面的未来表示,结合3D和语义模态,提升环境理解与现实推理能力,从而改进轨迹规划。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28196 2026-05-01 cs.CV 94%

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

HERMES++:迈向统一的驾驶世界模型用于3D场景理解和生成

Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan, Dingyuan Zhang, Hengshuang Zhao, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Mach Drive University of Hong Kong(香港大学)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 本文提出HERMES++,一种统一的驾驶世界模型,整合3D场景理解和未来几何预测。通过BEV表示、LLM增强世界查询和当前到未来链接等设计,提升驾驶场景的生成与理解能力。

Comments Extended version of ICCV 25 paper HERMES, Code: https://github.com/H-EmbodVis/HERMESV2, Project page: https://h-embodvis.github.io/HERMESV2/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14729 2025-08-14 cs.CV 94%

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

Xin Zhou, Dingkang Liang, Sifan Tu, Xiwu Chen, Yikang Ding, Dingyuan Zhang, Feiyang Tan, Hengshuang Zhao, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) MEGVII Technology(梅格维七科技) Mach Drive(马奇驱动) The University of Hong Kong(香港大学)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

Comments Accepted by ICCV 2025. The code is available at https://github.com/LMD0311/HERMES

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24004 2026-05-26 cs.AI cs.CV cs.LG cs.RO 94%

Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving

推理--想象--行动:基于世界模型的闭环LLM自动驾驶决策

Zhengqi Sun, Yiwen Sun, Boxuan Liu, Tailai Chen, Tianxu Guo, Jiabin Liu

机构 * 1Department of Information Management, Peking University, Beijing 100871, China 2School of Intelligence Science Technology, Peking University, Beijing 100871, China 3State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing 100080, China 4Yuanpei College, Peking University, Beijing 100871, China 5China Agricultural University, Beijing, China 6CRSC Research \& Design Institute Group Co., Ltd., Beijing, China

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出Reason--Imagine--Act (RIA)闭环框架,结合LLM推理器与动作条件世界模型进行在线安全验证,在CARLA点目标协议下实现80.05%路线完成率、51.10%到达率和0.20%碰撞率。

Comments Accepted by the 2026 IEEE International Conference on Intelligent Transportation Systems (ITSC 2026). 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13840 2026-08-20 cs.RO cs.CV 版本更新 94%

Multi-Agent Embodied Autonomous Driving (MAEAD): From V2X Information Exchange to Shared World Models

多智能体具身自动驾驶:从V2X信息交换到共享世界模型

Senkang Hu, Zhengru Fang, Yihang Tao, Zihan Fang, Yiqin Deng, Yuguang Fang

机构 * Lingnan University, Hong Kong(岭南大学(香港))

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文综述了从单车智能向多智能体具身系统转变的自动驾驶技术,通过共享世界模型实现感知共享、意图推断和协同规划,并指出了在仿真评估、实时安全保证等方面的研究空白。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10386 2026-08-12 cs.LG cs.RO 新提交 94%

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

Dreamer-SAC:用于样本高效自动驾驶的潜世界模型离线策略学习

Jiazhuo Li, Linjiang Cao, Qi Liu, Xi Xiong

机构 * Tongji University(同济大学)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出Dreamer-SAC框架,结合循环状态空间世界模型与离线策略SAC算法,在自动驾驶场景中优于DreamerV3、SAC等基线,且所需真实环境交互更少。

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14005 2026-07-16 cs.CV cs.RO 新提交 94%

M$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

M$^\text{4}$World:用于交互式对象操纵和分钟级流的多视图多模态驾驶世界模型

Ke Cheng, Hanqiao Ye, Lei Shi, Yahui Liu, Yunhan Shen, Jingtao Dong, Zhenke Wang, Wenxuan Ao, Weixiang Xu, Kaining Huang, Shuhan Shen

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 针对现有驾驶世界生成方法局限,提出M$^\text{4}$World模型,通过灵活接口与多阶段训练实现对象操纵及长时流稳定,引入后训练与生成模型,并用新管道评估,实验证明其在驾驶模拟中有高质量、可控性与稳定性。

Comments 24 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01536 2026-02-03 cs.RO cs.CV 94%

UniDWM: Towards a Unified Driving World Model via Multifaceted Representation Learning

UniDWM: 通过多维表征学习实现统一的驾驶世界模型

Shuai Liu, Siheng Ren, Xiaoyao Zhu, Quanmin Liang, Zefeng Li, Qiang Li, Xin Hu, Kai Huang

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家) Xpeng Motors Technology Co Ltd(小鹏汽车科技有限公司)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 UniDWM通过多维表征学习实现统一驾驶世界模型,提升自动驾驶中的轨迹规划和4D重建能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15341 2026-06-16 cs.CV 新提交 93%

CausalDrive: Real-time Causal World Models for Autonomous Driving

CausalDrive: 用于自动驾驶的实时因果世界模型

Tianyi Yan, Huan Zheng, Dubing Chen, Meizhi Qu, Yingying Shen, Lijun Zhou, Mingfei Tu, Bing Wang, Guang Chen, Hangjun Ye, Haiyang Sun, Cheng-zhong Xu, Jianbing Shen

机构 * SKL-IOTSC, CIS, University of Macau(澳门大学协同创新研究院,科技学院) Xiaomi EV(小米汽车) CASIA(中国科学院自动化研究所)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出CausalDrive,一种可控、实时的驾驶世界渲染器,通过因果预测和Context-Forced DMD架构实现交互式模拟,支持闭环评估、强化学习后训练和人在环仿真。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09701 2026-05-12 cs.CV 93%

DriveFuture: Future-Aware Latent World Models for Autonomous Driving

DriveFuture: 用于自动驾驶的面向未来的潜在世界模型

Yufeng Hong, Xiaotian Zhou, Yingyan Li, Xiangpo Zhou, Lin Liu, Yadan Luo, Shaoqing Xu, Lei Yang, Ziying Song

机构 * Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beihang University(北航) Beijing Jiaotong University(北京交通大学) The University of Queensland(昆士兰大学) University of Macau(澳门大学) Nanyang Technological University(南洋理工大学) School of Artificial Intelligence ( School of Software), Yanshan University(燕山大学人工智能学院(软件学院))

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出DriveFuture,一种面向未来的潜在世界建模框架,通过将未来世界状态条件化于当前潜在状态建模过程,提升自动驾驶轨迹规划性能。

Comments 24pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00969 2026-04-02 cs.CV 93%

DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving

DLWM:双潜在世界模型实现自动驾驶中的整体高斯中心预训练

Yiyao Zhu, Ying Xue, Haiming Zhang, Guangfeng Jiang, Wending Zhou, Xu Yan, Jiantao Gao, Yingjie Cai, Bingbing Liu, Zhen Li, Shaojie Shen

机构 * HKUST(香港科技大学) CUHK-SZ(香港中文大学(深圳)) USTC(中国科学技术大学) Huawei Foundation Model Department(华为基础模型部门)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出DLWM,通过双潜在世界模型实现自动驾驶中的整体高斯中心预训练,提升3D占用感知、4D占用预测和运动规划性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12751 2025-12-16 cs.CV 93%

GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation

GenieDrive: 向具有物理意识的驾驶世界模型迈进:基于4D占用的视频生成

Zhenya Yang, Zhe Liu, Yuxiang Lu, Liping Hou, Chenxuan Miao, Siyi Peng, Bailan Feng, Xiang Bai, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Huazhong University of Science and Technology(华中科技大学)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 GenieDrive通过4D占用引导的视频生成,实现物理意识的驾驶视频生成,提升预测精度和视频质量。

Comments The project page is available at https://huster-yzy.github.io/geniedrive_project_page/

详情

展开后加载摘要…

URL PDF HTML 收藏