arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-05-01 至 2026-05-01 共收录 14 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 8 篇

2604.27895 2026-05-01 cs.AI 93%

Graph World Models: Concepts, Taxonomy, and Future Directions

图世界模型:概念、分类与未来方向

Jiawei Liu, Senqiao Yang, Mingjun Wang, Yu Wang, Bei Yu

机构 * The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文系统阐述了图世界模型的概念,分类了基于关系归纳偏置的三种类型,并探讨了其未来研究方向与挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27935 2026-05-01 cs.RO cs.SY eess.SP eess.SY 92%

Flying by Inference: Active Inference World Models for Adaptive UAV Swarms

通过推断飞行:用于自适应无人机群的主动推断世界模型

Kaleem Arshid, Ali Krayani, Lucio Marcenaro, David Martin Gomez, Carlo Regazzoni

机构 * Department of Engineering and Naval Architecture (DITEN), University of Genoa(工程与 naval 架构系(DITEN),热那亚大学) Intelligent Systems Laboratory, Department of Systems Engineering and Automation, Carlos III University of Madrid(智能系统实验室,系统工程与自动化系,马德里卡洛斯三世大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文提出了一种专家引导的主动推断框架,用于自适应无人机群轨迹规划。该方法将多无人机轨迹设计转化为分层概率推断问题,通过遗传算法生成专家示范并学习世界模型,实现高效的轨迹规划与动态调整。

Comments Submitted to IEEE journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24329 2026-05-01 cs.CL 90%

World model inspired sarcasm reasoning with large language model agents

受世界模型启发的大型语言模型代理 sarcasm 推理

Keito Inoshita, Shinnosuke Mizuno

机构 * Faculty of Business and Commerce, Kansai University(大阪 kansai 大学 商业与文理学院) Faculty of Medicine, The University of Tokyo(东京大学 医学部)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract)

AI总结 本文提出 WM-SAR 模型,通过分解语义不一致性和意图进行 sarcasm 推理,实现高可解释性和性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28122 2026-05-01 cs.CV cs.LG 82%

Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces

超越高斯瓶颈:基于拓扑对齐的视觉Transformer特征空间编码

Andrew Bond, Ilkin Umut Melanlioglu, Erkut Erdem, Aykut Erdem

机构 * Department of Computer Engineering, Koç University, Istanbul, Turkey(科克大学计算机工程系,伊斯坦布尔,土耳其) Department of Computer Engineering, Hacettepe University, Ankara, Turkey(哈恰塔佩大学计算机工程系,安卡拉,土耳其) KUIS AI Research Center, Istanbul, Turkey(KUIS人工智能研究中心,伊斯坦布尔,土耳其) Department of Electrical and Electronics Engineering, Koç University, Istanbul, Turkey(科克大学电气与电子工程系,伊斯坦布尔,土耳其)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出S²VAE框架,通过压缩和表示场景的3D状态,包括相机运动、深度和点结构,以提升视觉模型的几何一致性。实验显示,几何对齐的超球面隐空间在高压缩条件下优于传统高斯瓶颈。

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27576 2026-05-01 cs.LO cs.LG 81%

BAss: Symbolic Reasoning in Abstract Dialectical Frameworks

BAss:基于BDD的抽象辩证框架符号推理

Samuel Pastva, Van-Giang Trinh

机构 * Faculty of Informatics, Masaryk University(马萨里克大学信息学院) Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT)(胡志明市技术大学计算机科学与工程学院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 BAss基于BDD提出新型分析工具,实现抽象辩证框架的所有可接受、完整和优先解释的全符号计算,优于现有工具并在大规模解空间场景中表现优异,推动系统生物学研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04978 2026-05-01 cs.AI 81%

Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI

对齐感知、推理、建模与交互:物理AI的综述

Kun Xiang, Terry Jingchen Zhang, Yinya Huang, Jixi He, Zirong Liu, Yueling Tang, Ruizhe Zhou, Lijing Luo, Youpeng Wen, Xiuwei Chen, Bingqian Lin, Jianhua Han, Hang Xu, Hanhui Li, Bin Dong, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) ETH Zurich(苏黎世联邦理工学院) ETH AI Center(苏黎世联邦理工学院人工智能中心) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Yinwang Intelligent Technology Co., Ltd.(云网智能技术有限公司) Peking University and Beijing International Center for Mathematical Research(北京大学和北京国际数学研究中心)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文综述了物理AI,探讨了理论物理推理与应用物理理解的区别,并分析了基于物理的方法如何提升AI在现实世界中的理解能力,推动更安全、可推广的AI系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27955 2026-05-01 cs.AI cs.CV 71%

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants

具有强化学习的图形用户界面代理:迈向数字居民

Junan Hu, Jian Liu, Jingxiang Lai, Jiarui Hu, Yiwei Sheng, Shuang Chen, Jian Li, Dazhao Du, Song Guo

机构 * Shandong University(山东大学) The Hong Kong University of Science and Technology(香港科技大学) The Hong Kong University(香港大学) Shanghai Jiao Tong University(上海交通大学) Tencent(腾讯)

专题命中 通用世界模型 :world-model(abstract);world-model(abstract);分类 cs.AI、cs.CV

AI总结 本文探讨了强化学习与GUI代理的结合,提出分类体系及趋势分析,旨在指导下一代鲁棒GUI自动化及基础设施发展。

Comments Project Page: https://github.com/Steve2457/Awesome-RL-GUI-Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27899 2026-05-01 cs.AI 69%

Simulating clinical interventions with a generative multimodal model of human physiology

用生成式多模态模型模拟临床干预

Guy Lutsker, Gal Sapir, Jordi Merino, Smadar Shilo, Anastasia Godneva, Eli Meirom, Shie Mannor, Hagai Rossman, Gal Chechik, Eran Segal

机构 * Department of Computer Science and Applied Mathematics, Weizmann Institute of Science(魏茨曼科学研究所计算机科学与应用数学系) Department of Molecular Cell Biology, Weizmann Institute of Science(魏茨曼科学研究所分子细胞生物学系) NVIDIA Novo Nordisk Foundation Center for Basic Metabolic Research, University of Copenhagen(诺沃维克基金会基础代谢研究中心,哥本哈根大学) Faculty of Medical and Health Sciences, Tel Aviv University(特拉维夫大学医学与健康科学学院) The Jesse Z and Sara Lea Shafer Institute for Endocrinology and Diabetes, National Center for Childhood Diabetes, Schneider Children’s Medical Center of Israel(杰西Z和索菲亚·李·沙弗内分泌学与糖尿病研究所,以色列儿童糖尿病国家中心,施耐德儿童医学中心) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 本文提出HealthFormer模型,通过训练人类表型项目数据,生成人类生理轨迹,实现对个体生理变化的预测和干预模拟,提升临床风险评分和疾病预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 自动驾驶 1 篇

2604.28196 2026-05-01 cs.CV 94%

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

HERMES++:迈向统一的驾驶世界模型用于3D场景理解和生成

Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan, Dingyuan Zhang, Hengshuang Zhao, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Mach Drive University of Hong Kong(香港大学)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 本文提出HERMES++,一种统一的驾驶世界模型,整合3D场景理解和未来几何预测。通过BEV表示、LLM增强世界查询和当前到未来链接等设计,提升驾驶场景的生成与理解能力。

Comments Extended version of ICCV 25 paper HERMES, Code: https://github.com/H-EmbodVis/HERMESV2, Project page: https://h-embodvis.github.io/HERMESV2/

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 模型式强化学习 5 篇

2604.27450 2026-05-01 cs.RO cs.AI 77%

RAY-TOLD: Ray-Based Latent Dynamics for Dense Dynamic Obstacle Avoidance with TDMPC

RAY-TOLD: 基于射线的任务导向潜在动力学用于密集动态障碍物避障与TDMPC

Seungho Han, Seokju Lee, Jeonguk Kang

机构 * School of Electrical Engineering, Hanyang University(翰阳大学电气工程学院) Mechatronics, Systems and Control Lab (MSC Lab), Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(机械工程系,韩国科学技术院(KAIST)机电系统与控制实验室(MSC实验室)) Samsung Research, Samsung Electronics(三星研究所,三星电子)

专题命中 模型式强化学习 :latent dynamics(title,abstract);分类 cs.AI、cs.RO;dynamics model(abstract)

AI总结 本文提出RAY-TOLD,结合物理基础MPPI的鲁棒性与强化学习的长视界,通过LiDAR中心的潜在动力学模型实现动态障碍物避障,提升导航可靠性与安全性。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27411 2026-05-01 cs.LG 74%

Detecting is Easy, Adapting is Hard: Local Expert Growth for Visual Model-Based Reinforcement Learning under Distribution Shift

检测容易,适应困难:基于视觉模型的强化学习在分布偏移下的局部专家增长

Haiyang Zhao

机构 * University of Georgia(佐治亚大学)

专题命中 模型式强化学习 :model-based reinforcement learning(title,abstract);分类 cs.LG

AI总结 本文研究了视觉模型强化学习在分布偏移下的适应问题,提出JEPA-Indexed Local Expert Growth方法,通过局部专家进行动作修正,提升了对偏移环境的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24265 2026-05-01 cs.MA 58%

R3DM: Enabling Role Discovery and Diversity Through Dynamics Models in Multi-agent Reinforcement Learning

R3DM:通过动态模型在多智能体强化学习中实现角色发现与多样性

Harsh Goel, Mohammad Omama, Behdad Chalaki, Vaishnav Tadiparthi, Ehsan Moradi Pari, Sandeep Chinchali

专题命中 模型式强化学习 :dynamics model(title,abstract);分类 cs.MA

AI总结 R3DM通过动态模型最大化智能体角色、观察轨迹与预期未来行为间的互信息,提升多智能体协作效率,实验表明其在SMAC和SMACv2环境中显著提升胜率。

Comments 21 pages, To appear in the International Conference of Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28161 2026-05-01 cs.RO 56%

RopeDreamer: A Kinematic Recurrent State Space Model for Dynamics of Flexible Deformable Linear Objects

RopeDreamer:一种用于柔性可变形线性物体动态的运动学递归状态空间模型

Tim Missal, Lucas Domingues, Berk Guler, Simon Manschitz, Jan Peters, Paula Dornhofer Paro Costa

机构 * Technical University of Darmstadt(德意志技术大学) School of Electrical and Computer Engineering, Universidade Estadual de Campinas (UNICAMP)(坎皮纳斯州立大学电气与计算机工程学院) Instituto de Pesquisas Eldorado(Eldorado研究所) Honda Research Institute Europe GmbH(本田欧洲研究院) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Robotics Institute Germany (RIG)(德国机器人研究所) Centre for Cognitive Science(认知科学研究中心) Artificial Ingelligence Lab, Recod.ai(Recod.ai人工智能实验室)

专题命中 模型式强化学习 :latent dynamics(abstract);分类 cs.RO;dynamics model(abstract)

AI总结 本文提出结合递归状态空间模型与四元数运动链表示的潜变量框架,用于预测柔性可变形线性物体的状态,通过约束物理有效流形减少自交和非物理变形,提升长周期预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27313 2026-05-01 cs.LG cs.CV 56%

PINN-Cast: Exploring the Role of Continuous-Depth NODE in Transformers and Physics Informed Loss as Soft Physical Constraints in Short-term Weather Forecasting

PINN-Cast:探索连续深度NODE在Transformer中的作用及物理信息损失作为短期天气预报中的软物理约束

Hira Saleem, Flora Salim, Cormac Purcell

机构 * University of New South Wales(新南威尔士大学)

专题命中 模型式强化学习 :latent dynamics(abstract);分类 cs.LG、cs.CV

AI总结 本文提出连续深度Transformer编码器,结合Neural ODE动态和物理信息损失,提升短期天气预报的准确性与物理一致性。

Comments 14 pages, 4 Figures, Accepted in 26th International Conference on Computational Science (ICCS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏