arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-02-04 至 2026-02-04 共收录 10 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 10 篇

2602.03793 2026-02-04 cs.RO cs.CV 84%

BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks

BridgeV2W: 通过具身遮罩将视频生成模型与具身世界模型 bridging

Yixiang Chen, Peiyan Li, Jiabing Yang, Keji He, Xiangnan Wu, Yuan Xu, Kai Wang, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(模式识别新实验室(NLPR),自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Shandong University(山东大学)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);分类 cs.RO、cs.CV

AI总结 BridgeV2W 通过具身遮罩将视频生成模型与具身世界模型结合,解决坐标空间动作与像素空间视频的不匹配问题,并提升视频生成质量与跨具身统一性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13696 2026-02-04 cs.AI 83%

Building spatial world models from sparse transitional episodic memories

从稀疏的过渡性片段记忆中构建空间世界模型

Zizhan He, Maxime Daigle, Pouya Bashivan

机构 * Department of Computer Science, McGill University(麦吉尔大学计算机科学系) Department of Physiology, McGill University(麦吉尔大学生理学系) Mila, Université de Montréal(蒙特利尔大学Mila)

专题命中 具身推理 :world model(title,abstract);navigation(abstract);分类 cs.AI

AI总结 本文提出ESWM模型,通过稀疏片段记忆构建空间地图,实现环境探索和导航的高效适应能力。

Comments Accepted ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04075 2026-02-04 cs.LG cs.AI cs.CV 82%

Accurate and Efficient World Modeling with Masked Latent Transformers

基于掩码潜在变换器的精准高效世界建模

Maxime Burchi, Radu Timofte

机构 * Computer Vision Lab, CAIDAS \& IFI, University of Würzburg, Germany

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV、cs.LG

AI总结 EMERALD通过高效掩码潜在变换器实现精准高效的世界建模,首次在1000万环境步骤内超越人类专家性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03569 2026-02-04 cs.AI cs.LG 81%

EHRWorld: A Patient-Centric Medical World Model for Long-Horizon Clinical Trajectories

EHRWorld: 一个以患者为中心的医疗世界模型用于长周期临床轨迹

Linjie Mu, Zhongzhen Huang, Yannian Gu, Shengqian Qin, Shaoting Zhang, Xiaofan Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 EHRWorld通过因果序列范式训练,有效解决长周期临床模拟中的误差累积问题,提升医疗世界模型的稳定性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02393 2026-02-04 cs.CV cs.AI 81%

Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory

无限世界:通过无姿态分层记忆实现交互世界模型的1000帧扩展

Ruiqi Wu, Xuanhua He, Meng Cheng, Tianyu Yang, Yong Zhang, Zhuoliang Kang, Xunliang Cai, Xiaoming Wei, Chunle Guo, Chongyi Li, Ming-Ming Cheng

机构 * VCIP, CS, Nankai University(中国南开大学计算机科学与技术学院) The Hong Kong University of Science(香港科学与技术大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 Infinite-World通过无姿态分层记忆压缩和不确定性感知动作标签,实现交互世界模型在1000帧内的稳健扩展,提升视觉质量和动作可控性。

Comments project page: https://rq-wu.github.io/projects/infinite-world/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03747 2026-02-04 cs.CV 79%

LIVE: Long-horizon Interactive Video World Modeling

LIVE: 长时距交互视频世界建模

Junchao Huang, Ziyang Ye, Xinting Hu, Tianyu He, Guiyu Zhang, Shaoshuai Shi, Jiang Bian, Li Jiang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Loop Area Institute(深圳河套学院) Microsoft Research(微软研究院) The University of Hong Kong(香港大学) Voyager Research, Didi Chuxing Project(Voyager研究,滴滴出行项目)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 LIVE通过循环一致性目标限制误差累积,无需教师蒸馏,实现长时距交互视频世界建模,取得最优性能。

Comments 18 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03242 2026-02-04 cs.CV 79%

InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation

InstaDrive: 为真实和一致的视频生成实例感知的驾驶世界模型

Zhuoran Yang, Xi Guo, Chenjing Ding, Chiyu Wang, Wei Wu, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) SenseAuto(感etime) Tsinghua University(清华大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 InstaDrive通过实例感知机制提升驾驶视频生成质量,增强自动驾驶任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03146 2026-02-04 cs.AI 74%

General Agents Contain World Models, even under Partial Observability and Stochasticity

通用智能体包含世界模型,即使在部分可观测性和随机性下

Santiago Cifuentes

机构 * Dovetail Research(Dovetail研究)

专题命中 具身推理 :world model(title);分类 cs.AI

AI总结 本研究证明了即使在部分可观测性和随机性下,通用智能体仍需学习环境模型。

Comments 19 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02857 2026-02-04 cs.RO 72%

Latent Perspective-Taking via a Schrödinger Bridge in Influence-Augmented Local Models

通过影响增强的局部模型实现潜在的视角转换

Kevin Alcedo, Pedro U. Lima, Rachid Alami

机构 * Institute for Systems and Robotics(系统与机器人研究所) Universidade de Lisboa(里斯本大学) LAAS-CNRS(拉沙尔国家研究中心) Artificial and Natural Intelligence Toulouse Institute (ANITI)(图卢兹人工智能与自然智能研究所)

专题命中 具身推理 :world model(abstract,comments);navigation(abstract);分类 cs.RO

AI总结 本文提出了一种基于影响增强的局部模型和施罗德桥的神经符号框架,用于实现机器人在社交交互中的心理状态推断和决策规划。

Comments Extended Abstract & Poster, Presented at World Modeling Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17873 2026-02-04 cs.CV cs.AI 62%

SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model

SurgVidLM:迈向多粒度外科视频理解的大型语言模型

Guankun Wang, Junyi Wang, Wenjin Mo, Long Bai, Kun Yuan, Ming Hu, Jinlin Wu, Junjun He, Yiming Huang, Nicolas Padoy, Zhen Lei, Hongbin Liu, Nassir Navab, Hongliang Ren

机构 * The Chinese University of Hong Kong(香港中文大学) Sun Yat-sen University(中山大学) University of Strasbourg(斯特拉斯堡大学) Technical University of Munich(慕尼黑技术大学) Monash University(墨尔本大学) Centre for Artificial Intelligence and Robotics, HKISI-CAS(人工智能与机器人中心,HKISI-CAS) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 具身推理 :robotic(abstract);分类 cs.AI、cs.CV

AI总结 SurgVidLM通过多粒度分析提升外科视频理解能力,结合全局与局部机制实现更精确的手术流程解析。

详情

展开后加载摘要…

URL PDF HTML 收藏