arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-08-25 至 2026-08-25 共收录 36 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 27 篇

2608.22294 2026-08-25 cs.RO 新提交 94%

Beyond Instance Slots: Semantically Rich World Models for Physical Interaction Planning

超越实例槽:面向物理交互规划的语义丰富世界模型

Juntao Cheng, Jingkai Wang, Yijun Shen, Xiansheng Chen, Zhiwei Yu

机构 * Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究针对物理交互规划中实例槽无法明确实体任务角色的问题,提出SR-WM模型,通过功能角色绑定实现语义接口,在LIBERO模拟套件等评估中验证了其视觉动力学与规划决策的连接能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22421 2026-08-25 cs.AI 新提交 94%

Where World Models Break: Natural-Input Failure Discovery

世界模型何时失效:自然输入的故障发现

Zhanpeng Shi, Zi Liang, Rong Feng, Shiqin Tang, Xuyang Chen, Hongzong Li

机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院) The Hong Kong Polytechnic University(香港理工大学) City University of Hong Kong(香港城市大学) Department of Electrical and Computer Engineering, National University of Singapore(新加坡国立大学电气与计算机工程系) Generative AI Research and Development Center, The Hong Kong University of Science and Technology(香港科技大学生成式人工智能研发中心)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文针对现有世界模型评估忽略灾难性故障风险的问题,提出 BasinLens 方法,可在有限查询预算下发现自然输入引发的可复现、局部持续的世界模型故障,暴露常规基准未覆盖的漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23070 2026-08-25 cs.AI cs.CV 新提交 94%

From Generation to Simulation: How Far Are World Models from Being True Simulators?

从生成到模拟:世界模型距离真正的模拟器还有多远?

Tong Wang, Huan Deng, Mucheng Yang, Yang He, Xiaohui Kuang, Gang Zhao

机构 * Institute of Systems Engineering, Academy of Military Sciences(军事科学院系统工程研究院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究通过系统评估200篇相关成果,分析了潜在动力学等3条技术路线的世界模型在传统模拟器8项能力上的表现,指出其在状态反馈等方面的缺陷并提出6个研究方向。

Comments 42 pages, 23 figures, 2 tables. Project page: this https URL (https://github.com/AtongWang/world-model-simulators)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28455 2026-08-25 cs.RO cs.AI cs.LG 版本更新 94%

Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Structure in Passive Object-State World Models

被动对象状态世界模型中运动学、接触和物体恒存场的事件条件诊断

Yang Liu, Yuming Chen

机构 * College of Intelligent Robitcs and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出一种诊断协议,测试被动对象状态世界模型中的隐式物理场是否按事件类型组织,并验证场对齐表示对预测的功能影响。

Comments Revised version: updated the title and terminology to use a more conservative latent-structure framing, added a third independent training seed across the evaluated model architectures, and regenerated the corresponding analyses, figures, and numerical results. The public reproducibility repository has also been updated. The main conclusions remain unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23248 2026-08-25 cs.CL cs.AI 新提交 94%

Future Querying: Can LLMs Serve as Implicit Medical World Models?

未来查询:大语言模型能否作为隐式医学世界模型?

Siri Willems, James Butterworth, Lore Goetschalckx, Peter Vrancx, Philippe Modard, Elke Giets, Ludovic Denoyer

机构 * imec, AI-labs(imec人工智能实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出未来查询范式,探究LLMs能否作为隐式医学世界模型,其框架可基于非结构化临床文档运行,经微调的小型开放权重模型性能接近专有系统,在相关数据集上验证了LLMs可捕捉临床动态。

Comments This paper is accepted at The 1st MICCAI Workshop on Medical World Models (MICCAI-2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17959 2026-08-25 cs.AI cs.LG 版本更新 94%

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

面向基于神经符号世界模型的零样本任务迁移

Isidoro Tamassia, Lennert De Smet, Giuseppe Marra

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出一种神经符号世界模型,通过解耦观测重构与奖励预测,实现无需额外环境交互的零样本任务迁移,其泛化能力优于纯神经方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23189 2026-08-25 cs.CV 新提交 93%

EchoWM: Open and Enterable Omnimodal World Models

EchoWM:开放且可进入的全模态世界模型

Songchun Zhang, Yaowei Li, Junhao Zhuang, Weiyang Jin, Haoyu Wang, Xin Lu, Yilang Sun, Shiyi Zhang, Haoran Li, Xiaoxiao Ma, Yuming Li, Yijun Liu, Yaofeng Su, Yanwen Ma, Haoyu Wu, Zihan Su, Yue Ma, Lvmin Zhang, Haoyang Huang, Zeyue Xue, Anyi Rao, Nan Duan

机构 * HKUST(香港科技大学) PKU(北京大学) Joy Future Academy, JD(京东探索研究院) HKU(香港大学) THU(清华大学) USTC(中国科学技术大学) FDU(复旦大学) Beihang University(北京航空航天大学) Stanford University(斯坦福大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 该研究提出EchoWM全模态世界模型,围绕相机意图组织交互,构建互补数据引擎并采用渐进式训练等方法,在公开基准上实现强轨迹跟随与高视觉质量,支持跨视角交互及长序列多模态同步生成。

Comments 42 pages, 24 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22750 2026-08-25 cs.LG 新提交 93%

MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models

MOSH-WM:面向以对象为中心的世界模型的基于掩码的软哈密顿动力学

Zhekai Wang, Haoxiang Huang, Xiang Liu, Zhikang Chen, Yueqing Sun, Qi Gu, Shiji Zhou, Miao Liu, Sen Cui

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究提出MOSH-WM模型,通过掩码支撑构建软哈密顿动力学,在OBJ3D、CLEVRER数据集的视频预测任务中,相比基线显著降低误差,且误差积累更慢。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22278 2026-08-25 cs.RO 新提交 92%

DreamMimic: Learning Visuomotor Whole-Body Loco-Manipulation via World Model

DreamMimic:通过世界模型学习视觉运动全身移动操作

Jie Yin, Xingyu Lai

机构 * Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world-model(abstract)

AI总结 DreamMimic 框架通过世界模型辅助蒸馏,将特权教师策略提炼为基于视觉的人形机器人控制器,引入 PCG 平衡引导与探索,在 OMOMO 和 BEHAVE 上提升了视觉移动操作性能。

Comments accepted to IROS2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22764 2026-08-25 cs.LG 新提交 92%

LpWM: A Case for Sparse Representations in World Models

LpWM:世界模型中稀疏表示的一个应用案例

Yilun Kuang, Yash Dagade, Quentin Le Lidec, Lucas Maes, Randall Balestriero, Yann LeCun

机构 * NYU(纽约大学) Duke University(杜克大学) Mila(米拉) Brown University(布朗大学)

专题命中 通用世界模型 :world model(title);world models(title);world model(title);world models(title)

AI总结 本文提出LpWM模型,以稀疏表示替代密集表示建模动作条件潜在动力学,在PushT任务上规划成功率优于密集模型,且能揭示可解释的动力学结构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07265 2026-08-25 math.OC 版本更新 91%

Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control

用于控制的学习型世界模型中的度量非崩溃:逼近理论、有限样本几何保证与确定性规划迁移

Alain Bensoussan, Minh-Nhat Phung, Minh-Binh Tran

专题命中 通用世界模型 :world model(title);world models(title);world model(title);world models(title)

AI总结 该研究为非线性确定性控制系统的学习型世界模型构建三部分数学理论,含逼近理论、有限样本几何保证及确定性规划迁移方法,通过数值实验验证了相关方法的有效性。

Comments Revised version prepared in response to the editorial assessment. The main manuscript is 32 pages; detailed mathematical derivations have been moved to the accompanying Supplementary Material. The principal results and contributions are strengthened and clarified

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23526 2026-08-25 cs.AI 新提交 91%

Correcting a learned physical invariant improves world-model rollouts

修正已学习的物理不变量可改进世界模型的滚动预测

Richard Bao

专题命中 通用世界模型 :world-model(title,comments);world-model(title,comments);world model(abstract);world models(abstract)

AI总结 本研究针对仅用单摆视频训练的DreamerV3,通过无标签搜索发现其学习到类能量不变量,投影隐态回初始水平集可降低保守模型滚动误差,揭示世界模型存在从像素学物理约束却在想象时违反的失效模式。

Comments 10 pages, 5 figures. Code at this https URL (https://github.com/Zarand3r/world-model-invariants)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18234 2026-08-25 cs.RO cs.AI cs.LG 版本更新 91%

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

GigaBrain-WBC-0.5:一种用于与环境交互的鲁棒全身控制的行为世界模型

Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI、cs.LG、cs.RO

AI总结 该研究提出首个用于人形机器人全身控制的行为世界模型GigaBrain-WBC-0.5,通过训练因果Transformer联合预测动作、状态与指令分布,在多场景下实现鲁棒控制,成功率优于现有基线。

Comments 20 pages, 8 figures, 4 tables. Technical report. Project page: this https URL (https://shepherd1226.github.io/gigabrain-wbc-0.5/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08600 2026-08-25 cs.CV cs.AI cs.LG 版本更新 91%

Population-Scalable Multi-Agent World Modeling

支持种群规模扩展的多智能体世界建模

Renjie Zhao, Yuxiang Wu, Mingyu Zhang, Jiaxin Li, Sisi Li, He Li, Yimin Sheng, Tianxi Tan, Zhenkai Zhang, Jiao Liang, Jianyi Zhu, Yong-Lu Li

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 针对多智能体世界模型的种群可扩展性问题,提出无需重训练即可扩展至任意智能体数量的Khora模型,通过解耦世界状态演化与渲染实现跨视图一致性,验证了泛化性并构建了实时交互系统。

Comments Technical report. Project page: this https URL (https://rhos.ai/research/khora). Online demo: this https URL (https://ophilus.ai/khora)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23565 2026-08-25 cs.AI 新提交 90%

ReWorld: An Interactive World Model with Long-Horizon Memory

ReWorld:具有长程记忆的交互式世界模型

Zhifei Chen, Luozhou Wang, Guibao Shen, Dongyu Yan, Shuai Yang, Tianshuo Xu, Yihua Du, Wei Wang, Tianyi Gui, Lianghua Huang, Yingcong Chen

机构 * HKUST(GZ)(香港科技大学(广州)) ATH, Alibaba(阿里巴巴 ATH 实验室)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 ReWorld通过训练分离、推理约束解决交互式世界模型的控制与记忆矛盾,结合混合注意力等技术,在三轴评估中优于6种模型,实现长程视频生成与记忆。

Comments 21 pages, 9 figures. Project page: this https URL (https://zhifeichen097.github.io/ReWorld/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22197 2026-08-25 cs.LG 新提交 90%

On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models

世界模型策略学习与模仿世界-动作模型之间的能力分离

Yang Yu

机构 * Nanjing University(南京大学)

专题命中 通用世界模型 :world-model(title,abstract);world-model(title,abstract);world model(abstract);world model(abstract)

AI总结 该研究对比了直接行为克隆策略等三类策略,明确世界-动作模型学习与直接行为克隆的能力差异,指出观测演示无法识别动作效果,干预可实现零遗憾值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21582 2026-08-25 cs.LG 新提交 90%

Reading the Room: Implicit Confusion Encoding in Recurrent World Model States

读取房间:循环世界模型状态中的隐式困惑编码

Donald Aadithiyan

机构 * University of Moratuwa(莫拉图瓦大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 该研究发现RSSM架构世界模型(如DreamerV3)的隐藏状态$h_t$含隐式困惑信号,经线性探针、编辑验证其因果性,该信号可在多数控制任务中泛化。

Comments Accepted for presentation at the GlobalSouthAI D&I Workshop at IJCAI 2026. Held in Bremen, Germany on the 17th August 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21439 2026-08-25 cs.CV 新提交 90%

WorldMind: Decoupled Game World Model for State-Aware NPC Behavior

WorldMind:用于状态感知NPC行为的解耦游戏世界模型

Zhiyang Deng, Boran Zhang, Danze Chen, Yeying Jin

机构 * Tencent(腾讯)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本研究针对现有游戏世界模型中NPC行为与视频生成绑定的问题,提出首个解耦框架WorldMind,构建BOSS-140K数据集,实验显示其NPC行为更优。

Comments Project page: this https URL (https://teawhite.cn/worldmind_projectpage/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18422 2026-08-25 cs.CV 版本更新 86%

Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control

生成现实:基于交互式视频生成的人本世界模拟

Linxi Xie, Lisong C. Sun, Ashley Neall, Tong Wu, Shengqu Cai, Gordon Wetzstein

机构 * Stanford University(斯坦福大学) NYU Shanghai(纽约大学上海分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文提出了一种基于交互式视频生成的人本世界模拟系统,通过结合头部和手部姿态控制,提升虚拟环境的交互性和用户控制感。

Comments Project page here: this https URL (https://codeysun.github.io/generated-reality)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23452 2026-08-25 cs.RO cs.AI cs.LG 新提交 83%

Reward-Free Continual Adaptation for Resilient Space Robots

面向弹性空间机器人的无奖励持续适应

Andrej Orsula, Miguel Olivares-Mendez, Carol Martinez

机构 * University of Luxembourg(卢森堡大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对空间机器人硬件退化导致传统控制策略失效的问题,提出无奖励持续学习框架,利用隐态世界模型使智能体无需奖励即可适应环境,在三类模拟任务中验证了方法有效性。

Comments Accepted for publication at the Third Conference on AI in and for Space (SPAICE 2026) | The source code is available at this https URL (https://github.com/AndrejOrsula/space_robotics_bench)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22642 2026-08-25 cs.LG cs.AI 新提交 82%

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

Mol-JEPA:用于分子的多模态联合嵌入预测架构

Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff

机构 * University of Tübingen(蒂宾根大学) Boehringer Ingelheim(勃林格殷格翰) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Brown University(布朗大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对分子基础模型的化学无效增强等局限,提出Mol-JEPA多模态框架,利用模态掩码融入生化上下文,在基准测试中展现出优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22855 2026-08-25 hep-th 新提交 80%

Bouncing shellworld embedded in charged AdS spacetime

嵌入带电AdS时空的弹跳壳世界

Karma P. Sherpa, Rishi Pokhrel, Indra K. P. Chettri, Tanay K. Dey

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 该研究探讨嵌入五维带电AdS时空的壳世界宇宙演化,发现其非奇异弹跳可化解柯西视界不稳定性,引入体中长弦后弹跳仍存在且规避该不稳定性。

Comments 13 pages and 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23383 2026-08-25 cs.CV 新提交 69%

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

面向持续故事与交互世界的长时序视听生成

Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang

机构 * Joy Future Academy, JD(京东探索研究院)

专题命中 通用世界模型 :world-model(abstract);world-model(abstract);分类 cs.CV

AI总结 研究针对长时序视听生成的需求,提出JoyAI-Echo-1.5系统,含长视频与世界模型变体,通过专用技术实现跨镜头一致性等性能,在相关基准上取得领先结果,为生成连贯内容提供基础。

Comments Project page: this https URL (https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22358 2026-08-25 cs.LG physics.geo-ph 新提交 69%

Tracing the Unlabeled Storm: Cross-Variable Transfer in a Lagrangian Atmospheric JEPA Framework

追踪未标记的风暴:拉格朗日大气JEPA框架中的跨变量迁移

K M Anirudh, S Sandeep, Hariprasad Kodamana

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 该研究提出跨变量代理学习方法,基于M-JEPA在无降水监督下预训练,实现了优于ECMWF集合的季风降水预测,为大气表示迁移提供诊断框架。

Comments 8 pages, 3 figures, plus supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21596 2026-08-25 cs.CV 版本更新 69%

$ϕ$-Scene: Physically Grounded Image-to-3D Scene Reconstruction

$\phi$-Scene: 物理驱动的图像到3D场景重建

Haodong Li, Lulu Shao, Haolin Lu, Yu Fu, Yen-Ru Chen, Seemandhar Jain, Manmohan Chandraker

机构 * University of California San Diego(加州大学圣地亚哥分校)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 提出$\phi$-Scene方法,将单图像3D场景重建视为拓扑驱动的物理组装过程,通过SDF优化和刚体仿真解决穿透与不稳定接触问题,在3D-Front数据集上取得最优性能。

Comments Project page: this https URL (https://phi-scene.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07399 2026-08-25 stat.ML cs.LG 版本更新 69%

Automatic, Debiased, and Invariant Counterfactual Generation under General Interventions

通用干预下的自动、去偏和不变反事实生成

Raphael C Kim, Jingsen Zhu, Ramin Zabih, Michele Santacatterina

机构 * Cornell Tech(康奈尔科技) Cornell University(康奈尔大学) Department of Biostatistics, Department of Population Health(生物统计学系、人口健康系) New York University Grossman School of Medicine(纽约大学格罗斯曼医学院)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 提出ADIGen框架,结合Riesz回归、因果不变性和正交统计学习,实现通用干预下反事实生成的自动、去偏和不变性,并提供过剩风险界。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22079 2026-08-25 q-bio.MN 新提交 64%

DigiPhen: a new paradigm for building predictive models of biological systems

DigiPhen:构建生物系统预测模型的新范式

H. Steven Wiley, Angela Cintolesi, Niaz Bahar Chowdhury, Jaydeep P Bardhan, Song Feng, Steven S. Andrews, Herbert M Sauro, Kristin E. Burnum-Johnson, Scott E. Baker, Douglas Mans

专题命中 通用世界模型 :predictive model(title,abstract);predictive models(title,abstract)

AI总结 DigiPhen平台是构建生物系统预测模型的新范式,通过集成工作流构建生物数字表征,可预测遗传与环境变化对细胞表型的影响,助力生物系统再工程。

Comments 18 pages, 1 figure and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身与机器人 4 篇

2606.01027 2026-08-25 cs.RO 版本更新 88%

$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation

$\tau_0$-WM:一种用于机器人操作的统一视频-动作世界模型

Pengfei Zhou, Shengcong Chen, Di Chen, Jiaxu Wang, Rongjun Jin, Bingwen Zhu, Yike Pan, Songen Gu, Kuanning Wang, Shufeng Nan, Xingyu Qiu, Chenhao Qiu, Pu Yang, Yunuo Cai, Jianxiong Gao, Yifan Li, Yanwei Fu, Xiangyu Yue, Zhi Chen, Jianlan Luo

机构 * Shanghai Innovation Institute(上海创新研究院) AGIBOT Finch

专题命中 具身与机器人 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 提出$\tau_0$-WM,一个统一视频-动作世界模型,通过共享视频扩散骨干集成策略学习、视频预测和动作评估,在长时域和精细操作任务上优于基线。

Comments Our project homepge: this https URL (https://tau0-wm.github.io)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22067 2026-08-25 cs.RO cs.AI cs.CV cs.LG 新提交 75%

Inferring Action from Future Latent State for Robotic Manipulation

从未来隐状态推断机器人操纵动作

Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Jie Cheng, Peilin Huang, Han Fu, Zhuo Li, Xiaoxue Ren

机构 * DeepLeap Research(DeepLeap研究院)

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG、cs.CV

AI总结 本文提出无需视频生成的机器人操纵模型DELE-w0.5,通过从捕获动作相关物理结果的未来隐状态推断动作,在4项长程操纵任务的480次试验中,其性能优于最强基线47.5和30.7个百分点,实现最优表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21554 2026-08-25 cs.RO cs.LG cs.MA 新提交 72%

Model-Based Reinforcement Learning for Heterogeneous Multi-Robot Task Assignment Under Distribution Shifts

分布偏移下异构多机器人任务分配的模型强化学习方法

Daniel Garces, Sara Castro, Adrian Haimovich, Byron Crowe, Stephanie Gil

机构 * John A. Paulson School Of Engineering And Applied Sciences, Harvard University(哈佛大学约翰·A·保尔森工程与应用科学学院) Harvard Medical School(哈佛医学院) Beth Israel Deaconess Medical Center(贝斯以色列女执事医疗中心) Stanford University School of Medicine(斯坦福大学医学院)

专题命中 具身与机器人 :model-based reinforcement learning(title);分类 cs.LG、cs.RO、cs.MA

AI总结 本文提出一种预测感知自适应滚动框架,用于分布偏移下异构多机器人任务分配,在医院护理任务案例中,该方法较多种基线缩短了等待时间,提升了服务覆盖与尾部延迟性能。

Comments 34 pages, 14 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏