arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-06-16 至 2026-06-16 共收录 39 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 31 篇

2510.19728 2026-06-16 cs.LG cs.AI 50%

Enabling Granular Subgroup Level Model Evaluations by Generating Synthetic Medical Time Series

通过生成合成医疗时间序列实现细粒度亚组级别模型评估

Mahmoud Ibrahim, Bart Elen, Chang Sun, Gökhan Ertaylan, Michel Dumontier

机构 * Institute of Data Science, Faculty of Science and Engineering, Maastricht University(数据科学研究所,科学与工程学院,马斯特里赫特大学) Department of Advanced Computing Sciences, Faculty of Science and Engineering, Maastricht University(先进计算科学系,科学与工程学院,马斯特里赫特大学) VITO(VITO研究院)

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 本文提出一种框架,利用合成ICU时间序列数据训练和评估预测模型,特别是在细粒度人口亚组中。引入Enhanced TimeAutoDiff,通过分布对齐惩罚增强潜在扩散目标,减少真实-合成与真实-真实评估差距,提升亚组模型评估的鲁棒性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身与机器人 2 篇

2606.14981 2026-06-16 cs.RO cs.AI cs.LG 新提交 73%

Inference-time Policy Steering via Vision and Touch

通过视觉和触觉进行推理时策略引导

Yilin Wu, Zilin Si, Zeynep Temel, Oliver Kroemer, Andrea Bajcsy

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG、cs.RO

AI总结 提出ViTaL框架,通过视觉采样验证和触觉引导扩散编辑的双层优化,在推理时引导机器人策略,显著提升接触丰富操作任务的成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15768 2026-06-16 cs.RO cs.AI 新提交 71%

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

LaWAM: 用于高效动力学感知机器人策略的潜在世界行动模型

Jialei Chen, Kai Wang, Kang Chen, Shuaihang Chen, Feng Gao, Wenhao Tang, Zhiyuan Li, Weilin Liu, Zhuyu Yao, Boxun Li, Yuanbo Xu, Chao Yu

机构 * Tsinghua University(清华大学) Jilin University(吉林大学) Nankai University(南开大学) Peking University(北京大学) Harbin Institute of Technology(哈尔滨工业大学) Zhongguancun Academy(中关村学院) Striding.AI Infinigence AI

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.AI、cs.RO

AI总结 提出LaWAM模型,通过潜在视觉子目标预测场景变化,实现动力学感知的机器人控制,在多个基准上达到最优或竞争性成功率,且推理延迟低。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 自动驾驶 3 篇

2606.15341 2026-06-16 cs.CV 新提交 93%

CausalDrive: Real-time Causal World Models for Autonomous Driving

CausalDrive: 用于自动驾驶的实时因果世界模型

Tianyi Yan, Huan Zheng, Dubing Chen, Meizhi Qu, Yingying Shen, Lijun Zhou, Mingfei Tu, Bing Wang, Guang Chen, Hangjun Ye, Haiyang Sun, Cheng-zhong Xu, Jianbing Shen

机构 * SKL-IOTSC, CIS, University of Macau(澳门大学协同创新研究院,科技学院) Xiaomi EV(小米汽车) CASIA(中国科学院自动化研究所)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出CausalDrive,一种可控、实时的驾驶世界渲染器,通过因果预测和Context-Forced DMD架构实现交互式模拟,支持闭环评估、强化学习后训练和人在环仿真。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16274 2026-06-16 cs.CV 新提交 92%

GraphWorld: Long-Horizon Planning with World Models for End-to-End Autonomous Driving

GraphWorld: 基于世界模型的长时域规划实现端到端自动驾驶

Ziying Song, Caiyan Jia, Lin Liu, Lei Yang, Shengkai Zhang, Feiyang Jia, Fengda Zhao, Peiliang Wu, Shaoqing Xu, Chen Lv, Yadan Luo

机构 * Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院,交通数据挖掘与具身智能北京市重点实验室) School of Artificial Intelligence (School of Software), Yanshan University(燕山大学人工智能学院(软件学院)) School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院) University of Macau(澳门大学) The University of Queensland(昆士兰大学)

专题命中 自动驾驶 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 提出GraphWorld框架,通过潜在世界建模增强长时域规划,利用自车中心交互图建模邻车关系,并基于世界状态条件规划实现安全轨迹生成,显著降低碰撞率。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12560 2026-06-16 cs.CV cs.LG cs.RO 版本更新 91%

CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving

CoIRL-AD:面向自动驾驶的潜在世界模型中的协作-竞争模仿-强化学习

Xiaoji Zheng, Ziyuan Yang, Yanhao Chen, Yuhang Peng, Yuanrong Tang, Gengyuan Liu, Bokui Chen, Jiangtao Gong

机构 * University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

专题命中 自动驾驶 :world model(title);world models(title);world model(title);world models(title)

AI总结 提出CoIRL-AD框架,通过解耦模仿学习与强化学习、利用潜在世界模型进行长时程奖励估计以及引入竞争机制,在离线训练中提升自动驾驶的鲁棒性,尤其在跨城市泛化和长尾场景中表现优异。

Comments 19 pages, 22 figures, ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 模型式强化学习 2 篇

2602.05352 2026-06-16 cs.LG math.SG 版本更新 53%

Smoothness Errors in Dynamics Models and How to Avoid Them

动力学模型中的平滑误差及如何避免

Edward Berman, Luisa Li, Jung Yeon Park, Robin Walters

专题命中 模型式强化学习 :dynamics model(title,abstract);分类 cs.LG

AI总结 本文研究了不同GNN在动力学建模中的平滑效应,证明了单位ary卷积对这类任务有害,并提出放松的单位ary卷积以平衡平滑性保留与物理系统需求。

Comments Ecstatic to share relaxed unitary mesh convolutions with the community :D! This version contains the camera ready for ICML 2026. Send me an email with your thoughts! I love getting mail :^)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16607 2026-06-16 eess.SP cs.IT cs.LG math.IT 新提交 50%

Context-Aware Markov VAE for CSI Compression in Wireless Systems

面向无线系统中CSI压缩的上下文感知马尔可夫VAE

Efstathios Chatziloizos, Konstantinos Vandikas, Aneta Vulgarakis Feljan, Zheng Chen, Nikolaos Pappas

机构 * Ericsson Research(爱立信研究) Knut and Alice Wallenberg Foundation(克努特和阿莉斯·瓦伦贝格基金会) ELLIIT Lund University(隆德大学)

专题命中 模型式强化学习 :latent dynamics(abstract);分类 cs.LG

AI总结 提出基于k-记忆马尔可夫变分自编码器的上下文感知压缩框架,利用有限时间窗口捕捉CSI在潜在空间中的演化,在低中压缩率下显著提升重构性能。

Comments 5 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 仿真与规划 1 篇

2606.16480 2026-06-16 cs.RO cs.AI cs.SY eess.SY 新提交 71%

HOLO-MPPI: Multi-Scenario Motion Planning via Hierarchical Policy Optimization

HOLO-MPPI:通过分层策略优化的多场景运动规划

Youngjae Min, Jovin D'sa, Faizan M. Tariq, David Isele, Navid Azizan, Sangjae Bae

机构 * Massachusetts Institute of Technology(麻省理工学院) Honda Research Institute, USA(本田研究所(美国))

专题命中 仿真与规划 :world model(abstract);world model(abstract);分类 cs.AI、cs.RO

AI总结 提出HOLO-MPPI框架,结合离线高层策略学习与在线低层随机最优控制,实现多场景运动规划,无需针对每个场景重新调整参数,在自动驾驶中优于MPPI和端到端RL基线。

详情

展开后加载摘要…

URL PDF HTML 收藏