arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-03-18 至 2026-03-18 共收录 14 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 10 篇

2603.16860 2026-03-18 cs.RO 96%

DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models

DreamPlan: 通过视频世界模型高效强化微调视觉语言规划器

Emily Yue-Ting Jia, Weiduo Yuan, Tianheng Shi, Vitor Guizilini, Jiageng Mao, Yue Wang

机构 * USC Physical Superintelligence Lab(USC物理超智能实验室) Toyota Research Institute(丰田研究院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 DreamPlan通过视频世界模型高效强化微调视觉语言规划器,利用零样本VLM生成探索数据训练动作条件视频生成模型,再通过ORPO在虚拟环境中微调VLM,提升物体 manipulation 成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13644 2026-03-18 cs.RO cs.AI cs.CV 94%

World Models for Learning Dexterous Hand-Object Interactions from Human Videos

为从人类视频学习灵巧手-物体交互构建世界模型

Raktim Gautam Goswami, Amir Bar, David Fan, Tsung-Yen Yang, Gaoyue Zhou, Prashanth Krishnamurthy, Michael Rabbat, Farshad Khorrami, Yann LeCun

机构 * FAIR at Meta(Meta 的 FAIR 部门) New York University(纽约大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出DexWM模型,通过手部关键点提取实现对精细手部动作的建模,提升未来状态预测和零样本迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01983 2026-03-18 q-fin.MF math.PR 91%

Real-world models for multiple term structures: a unifying HJM semimartingale framework

现实世界中多个期限结构的模型:一种统一的HJM半鞅框架

Claudio Fontana, Eckhard Platen, Stefan Tappe

专题命中 通用世界模型 :world model(title);world models(title);world model(title);world models(title)

AI总结 本文提出一种统一的HJM半鞅框架,用于建模金融、保险和能源市场中的多个期限结构,研究市场可行性并刻画局部鞅折价器集合,分析相关的随机偏微分方程(SPDE)的解的存在性和唯一性、不变性及仿射实现的存在性。

Comments 47 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14497 2026-03-18 cs.CV cs.RO 90%

WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning

WorldVLM:结合世界模型预测与视觉语言推理

Stefan Englmeier, Katharina Winter, Fabian B. Flohr

机构 * Munich University of Applied Sciences(慕尼黑应用科学大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 WorldVLM结合视觉语言模型与世界模型,通过统一架构提升自动驾驶中的环境预测与决策能力,解决空间理解受限问题。

Comments 8 pages, 6 figures, 5 tables; submitted to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16871 2026-03-18 cs.CV 81%

WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation

WorldCam: 交互式自回归3D游戏世界与相机姿态作为统一的几何表示

Jisu Nam, Yicong Hong, Chun-Hao Paul Huang, Feng Liu, JoungBin Lee, Jiyoung Kim, Siyoon Jin, Yunsung Lee, Jaeyoon Jung, Suhwan Choi, Seungryong Kim, Yang Zhou

机构 * Adobe Research(Adobe研究院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出WorldCam,通过将相机姿态作为统一的几何表示,解决交互式游戏世界中动作控制与长时序3D一致性问题,结合物理连续动作空间和全局相机姿态索引提升导航与视觉质量。

Comments Project page is available at https://cvlab-kaist.github.io/WorldCam/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16306 2026-03-18 cs.CV 69%

DriveFix: Spatio-Temporally Coherent Driving Scene Restoration

DriveFix:时空一致的驾驶场景修复

Heyu Si, Brandon James Denis, Muyang Sun, Dragos Datcu, Yaoru Li, Xin Jin, Ruiju Fu, Yuliia Tatarinova, Federico Landi, Jie Song, Mingli Song, Qi Guo

机构 * Zhejiang University(浙江大学) Huawei(华为)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 DriveFix提出一种多视角修复框架,通过交错扩散变换器架构和几何感知损失,实现驾驶场景时空一致性,提升4D世界建模的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16723 2026-03-18 cs.LG cs.AI 50%

Federated Learning with Multi-Partner OneFlorida+ Consortium Data for Predicting Major Postoperative Complications

基于多合作伙伴OneFlorida+联盟数据的联邦学习用于预测重大术后并发症

Yuanfang Ren, Varun Sai Vemuri, Zhenhong Hu, Benjamin Shickel, Ziyuan Guan, Tyler J. Loftus, Parisa Rashidi, Tezcan Ozrazgat-Baslanti, Azra Bihorac

机构 * Intelligent Clinical Care Center, University of Florida(佛罗里达大学智能临床护理中心) Department of Medicine, Division of Nephrology, Hypertension, and Renal Transplantation, University of Florida(佛罗里达大学医学系肾病、高血压和肾移植分部) Department of Surgery, University of Florida(佛罗里达大学外科系) Department of Biomedical Engineering, University of Florida(佛罗里达大学生物医学工程系)

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 本研究利用多中心数据开发并验证了联邦学习模型,以预测重大术后并发症和死亡率,展示了联邦学习在隐私保护下的强大泛化能力。

Comments 1 figure, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16377 2026-03-18 cs.LG cs.AI 50%

Age Predictors Through the Lens of Generalization, Bias Mitigation, and Interpretability: Reflections on Causal Implications

通过泛化、偏差缓解和可解释性视角的年龄预测器:对因果影响的反思

Debdas Paul, Elisa Ferrari, Irene Gravili, Alessandro Cellerino

机构 * Leibniz Institute on Aging — Fritz Lipmann Institute (FLI)(利普希茨衰老研究所——弗里茨·利普曼研究所) Scuola Normale Superiore(正常大学)

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 本文探讨了通过消除外源性属性如种族、性别或组织来提升年龄预测的泛化能力,讨论了可解释神经网络模型在因果推断中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16330 2026-03-18 cs.CV cs.AI cs.LG 50%

An Interpretable Machine Learning Framework for Non-Small Cell Lung Cancer Drug Response Analysis

一种可解释的机器学习框架用于非小细胞肺癌药物反应分析

Ann Rachel, Pranav M Pawar, Mithun Mukharjee, Raja M, Tojo Mathew

机构 * Department of Computer Science(计算机科学系) Birla Institute of Technology and Science Pilani(比拉理工学院和科学学院)

专题命中 通用世界模型 :分类 cs.AI、cs.LG、cs.CV;predictive model(abstract)

AI总结 本文提出基于多组学数据的可解释机器学习框架,利用XGBoost预测肺癌药物反应,并通过SHAP和DeepSeek解释关键基因和通路的作用。

Comments 26 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15651 2026-03-18 cs.LG cs.AI 50%

A federated learning framework with knowledge graph and temporal transformer for early sepsis prediction in multi-center ICUs

一种融合知识图谱和时间变换器的联邦学习框架用于多中心ICU早期脓毒症预测

Yue Chang, Guangsen Lin, Jyun Jie Chuang, Shunqi Liu, Xinkui Li, Yaozheng Li

机构 * Chengdu Medical College(成都医学院) Kunming Medical University(昆明医科大学) National Yang Ming Chiao Tung University(国立阳明交通大学) University of Southern California(南加州大学) Yanshan University(燕山大学)

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 本文提出融合知识图谱和时间变换器的联邦学习框架,用于多中心ICU早期脓毒症预测,通过隐私保护的协同训练提升预测性能,实验表明其AUC达到0.956,优于传统集中模型和标准联邦学习。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身与机器人 2 篇

2603.16669 2026-03-18 cs.RO cs.CV 85%

Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation

Kinema4D: 用于空间时间具身模拟的运动学4D世界建模

Mutian Xu, Tianbao Zhang, Tianqi Liu, Zhaoxi Chen, Xiaoguang Han, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S-Lab) SSE, CUHKSZ(香港中文大学深圳校区SSE)

专题命中 具身与机器人 :world model(title);world model(title);分类 cs.CV、cs.RO

AI总结 本文提出Kinema4D,一种基于动作条件的4D生成机器人模拟器,通过精确的4D机器人控制表示和环境反应的生成建模,实现了高精度的空间时间交互模拟,并首次展示零样本迁移能力。

Comments Project page: https://mutianxu.github.io/Kinema4D-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16086 2026-03-18 cs.RO cs.AI cs.CV cs.SD 73%

Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation

迈向视觉-声音-语言-动作范式:用于以声音为中心的操作的HEAR框架

Chang Nie, Tianchen Deng, Guangming Wang, Zhe Liu, Hesheng Wang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University and Shanghai Key Laboratory of Navigation and Location Based Services(自动化与智能感知学院,上海交通大学,导航与基于位置的服务重点实验室) Department of Engineering, Cambridge University(工程系,剑桥大学)

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.AI、cs.CV、cs.RO

AI总结 本文提出HEAR框架,通过整合声音、视觉、语言和本体感知,解决实时声音中心操作中的关键声音遗漏问题,强调因果持续性和显式时间学习的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 自动驾驶 1 篇

2603.15771 2026-03-18 cs.RO cs.AI 76%

CorrectionPlanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving

CorrectionPlanner:基于强化学习的自主驾驶自纠正规划器

Yihong Guo, Dongqiangzi Ye, Sijia Chen, Anqi Liu, Xianming Liu

机构 * Department of Computer Science, Johns Hopkins University, Baltimore, USA.(约翰霍普金斯大学计算机科学系)

专题命中 自动驾驶 :world model(abstract);world model(abstract);model-based reinforcement learning(abstract);分类 cs.AI、cs.RO

AI总结 本文提出CorrectionPlanner,通过自纠正机制在规划过程中生成安全动作,减少碰撞率,提升规划性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 模型式强化学习 1 篇

2603.15857 2026-03-18 cs.AI cs.LG cs.RO 77%

Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models

正则化潜在动态预测是行为基础模型的强基线

Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White

机构 * Department of Computing Science, University of Alberta, Canada(阿尔伯塔大学计算机科学系) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所) Canada CIFAR AI Chair(加拿大CIFAR人工智能 chair) The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)

专题命中 模型式强化学习 :latent dynamics(title,abstract);分类 cs.AI、cs.LG、cs.RO

AI总结 本文探讨零样本强化学习中复杂表征学习目标的必要性,提出正则化潜在动态预测方法,通过正则化保持特征多样性,优于现有方法,并在低覆盖场景中表现优异。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏