arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-09-01 至 2026-09-01 共收录 24 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 24 篇

2608.29029 2026-09-01 cs.LG cs.AI 新提交 94%

Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World Models

Flow-JEPA:用于JEPA世界模型中鲁棒潜在动力学的流匹配方法

Yanchen Huo, Ziying Song, Yadan Luo

机构 * Nanyang Technological University(南洋理工大学) The University of Queensland(昆士兰大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究针对JEPA世界模型确定性自回归预测器易累积误差且对扰动敏感的问题,提出Flow-JEPA,采用条件流匹配实现随机轨迹级预测,显著提升了干净及含噪观测下的平均成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29434 2026-09-01 cs.LG cs.AI cs.CV 新提交 94%

Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations

潜在规划在点云场景中是否可行?面向几何观测的动作条件JEPA世界模型

Fabio F. Oberweger, Michael Schwingshackl

机构 * AIT Austrian Institute of Technology(奥地利技术研究所) Assistive & Autonomous Systems(辅助与自主系统)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究将JEPA模型扩展至点云场景,验证了三种JEPA设计可在点云规划中有效工作,其中动作敏感型模型表现最优,还实现了无需目标观测的3D目标接口构建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23602 2026-09-01 cs.RO cs.AI cs.LG 版本更新 94%

Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models

物理空间中相邻集的动作优于世界模型中的最佳预测

Liangyu Li, Qingwen Liu, Mingqing Liu, Wen Fang

机构 * Tongji University(同济大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究基于采样和潜在世界模型的控制器存在的条件失败提议过度生成问题,提出相邻集动作重建(ASAR)方法,在携带和释放评估集上,相比匹配选择提高了事件完成成功率,还刻画了选择风险。

Comments 23 pages, 7 figures. Includes supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30067 2026-09-01 cs.LG cs.AI cs.CL 新提交 94%

How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account

世界模型与策略如何在大语言模型智能体中结合?一种联合谱分析与行为学解释

Ruize Xu, Xiao Yu, Yujin Tang, Chenming Shang, Nikhil Singh

机构 * Dartmouth College(达特茅斯学院) Columbia University(哥伦比亚大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 该研究通过受控实验探究LLM智能体中世界模型与策略的结合机制,从几何与行为角度揭示二者的互补特性,并提出提升策略训练保留世界知识能力的方法。

Comments Accepted to EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03400 2026-09-01 cs.LG cs.AI 版本更新 94%

Better World Models Can Lead to Better Post-Training Performance

更好的世界模型可以导致更好的训练后性能

Prakhar Gupta, Henry Conklin, Sarah-Jane Leslie, Andrew Lee

机构 * University of Michigan(密歇根大学) Princeton University(普林斯顿大学) Harvard University(哈佛大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本研究通过比较不同世界建模策略,发现显式建模能提升Transformer的状态表示质量,从而增强强化学习后训练的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29910 2026-09-01 cs.CV 新提交 93%

Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

Matrix-Game 3.5:利用补丁内存增强实时流式交互式世界模型

Runjia Qian, Zile Wang, Jihai Zhang, Kai Zou, Wei Yu, Jiaxing Li, Zexiang Liu, Yaokun Li, Fei Kang, Kaichen Huang, Mengyin An, Haobo Zhang, Biao Jiang, Jiahua Wang, Haofeng Sun, Yang Liu, Yangguang Li

机构 * Riemann Dynamics(黎曼动力学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 Matrix-Game 3.5通过三项关键改进优化交互式世界模型,实现稳定长时序实时交互式生成,在多类任务上表现出色。

Comments https://matrix-game-v3-5.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03603 2026-09-01 cs.CV cs.CL 版本更新 93%

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

世界模型遇见语言模型:论具体推理与抽象推理的互补性

Yucheng Zhou, Wei Tao, Yiwen Guo, Jianbing Shen

机构 * Nanyang Technological University(南洋理工大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出受控具体推理框架及PF-OPSD方法,通过结合世界模型的视觉模拟与多模态大语言模型的抽象推理,在空间前瞻和开放域物理预测任务上提升性能与鲁棒性。

Comments EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28541 2026-09-01 cs.LG cs.AI cs.SY eess.SY 交叉投稿 93%

An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models

闭合模式是一种规范选择:认证代码世界模型中相对于可达性的拓扑结构

Javier Aguilar Martín

机构 * AGILabs

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 该研究探讨认证代码世界模型中相对于可达性的拓扑,通过LLM合成实验得出三条原则,揭示危险与可达性的关联、修复的限制及缓解措施需匹配错误维度方向的结论。

Comments 33 pages, 2 figures. Paper 3 of a series (companion papers: arXiv:2607.14169, arXiv:2608.17956). Code, data, and Lean formalization: https://github.com/JaviMaligno/code-world-models

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27367 2026-09-01 cs.CV cs.AI 版本更新 93%

Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

连续容量增长:面向JEPA世界模型的视觉Transformer编码器的任务复杂度驱动的宽度与深度扩展

Frederik Berenz

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文提出SCG方法,使JEPA世界模型的视觉Transformer编码器可按需连续扩展宽度或深度,在多任务上提升性能并实现更高参数效率,且零误扩展。

Comments 12 pages, 2 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15156 2026-09-01 cs.RO cs.AI 版本更新 93%

Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models

用于学习世界模型中反事实滚动的低秩动力学有效潜在载体

Yang Liu, Yuming Chen

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文针对学习世界模型的反事实滚动问题,提出基于低秩动力学有效潜在载体的方法,在两物体碰撞环境中验证秩4修补块可实现稳定的12步自主反事实滚动,且效果可复现。

Comments Revised manuscript with expanded Joint-intervention, carrier-relative, and recurrent-dynamics analyses. 54 pages, 7 figures. Code and data are available at https://github.com/lysea8282/dynamic-effective-latent-carriers

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29998 2026-09-01 cs.LG 新提交 92%

The Intervention Gap in Latent World Models

潜在世界模型中的干预差距

Donna Vakalis

机构 * Mila – Quebec AI Institute(米拉-魁北克人工智能研究所) University of Montreal(蒙特利尔大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 该研究发现潜在世界模型存在干预差距,即模型开环转移与环境干预对任务变量的影响存在偏差,需以捕获优先的方式在模型原生接口直接审计干预保真度。

Comments 21 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29242 2026-09-01 cs.RO 新提交 92%

AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization

AnyWorld:用于跨 embodiments 泛化的分解式第一人称世界模型

Cheng Chen, Jerry Bai, Jiacheng Wei, Boyu Chen, Xiaoji Zheng, Fan Wu, Minghao Yang, Tianrun Chen, Ruibo Li, Xiaoyu Yue, Xiaoyang Guo, Yixiao Ge, Guosheng Lin, Fayao Liu

机构 * Nanyang Technological University(南洋理工大学) Institute of Advanced Intelligence and Computing, A*STAR(新加坡科技研究局高级智能与计算研究院) XPENG Robotics(小鹏机器人) Zhejiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 AnyWorld 框架通过分解交互为动作、相机、embodiment 因子,将人类交互重组为机器人原生经验,生成数据可提升机器人操作性能,验证了动作校准与视觉重组的必要性。

Comments Project page: https://xpeng-robotics.github.io/anyworld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12564 2026-09-01 cs.LG 版本更新 92%

Scaling Automatic Research Agents via World Models

通过世界模型扩展自动研究智能体

Xiyuan Yang, Sheikh Sarwar, Jingru Cheng, Zhan Shi, Duanshun Li, Huiyuan Chen, Haiyang Zhang, Xing Fan, Chenlei Guo, Jingrui He, Zhenyu Liao

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊公司)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 针对自动研究智能体扩展时的训练瓶颈,提出WMRL方法,结合两种缓解措施提升收敛性,训练加速3-4倍且性能优于更大规模智能体,还可迁移至多类任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30897 2026-09-01 cs.AI 新提交 92%

CAER: Causal Action Effect Reweighting for World Model Training

CAER:面向世界模型训练的因果动作效应重加权

Jianjie Fang, Xvyuan Liu, Ziyou Wang, Rongze Tang, Zhaolu Wang, Zhuohang Li, Xin Zhang, Haisheng Su, Chen Gao, Wei Wu, Xinlei Chen, Yong Li

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world-model(abstract)

AI总结 针对现有动作条件世界模型训练中交互动态未充分优化的问题,提出CAER范式,通过在线定位动作因果影响的令牌并重加权,提升了生成视频的物理一致性、可控性与视觉质量。

Comments 14 pages, 8 figures. Project page: https://manifoldai-research.github.io/CAER/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30237 2026-09-01 cs.RO cs.AI cs.CV cs.LG 新提交 91%

Motus2: A Self-Evolving General World Model for Dexterous Manipulation

Motus2:一种用于灵巧操作的自进化通用世界模型

Hongzhe Bi, Zihao Zhou, Yihang Tang, Jingrui Pang, Shuhe Huang, Haitian Liu, Runqing Wang, Shuai Huang, Yichen Wang, Yiming Cheng, Ruowen Zhao, Zhenghua Li, Hengkai Tan, Xiaolong Liu, Jinhui Wan, Jiabao Liu, Min Zhao, Fan Bao, Jun Zhu

机构 * GensPI Tsinghua University(清华大学) BUAA(北京航空航天大学) BIT(北京理工大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出Motus2,一种用于灵巧操作的自进化通用世界模型,通过模型缩放与数据缩放构建闭环决策学习循环,结合多模态数据与仿生平台实现通用具身智能体的灵巧操作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28712 2026-09-01 cs.RO cs.LG 版本更新 91%

J-LAW: Joint Localization and Action-Conditioned World Modeling via Coupled Latent Factor Graphs

J-LAW:通过耦合潜在因子图实现联合定位与可操作世界建模

Guanqun Cao, Liang Chen

机构 * Geely Technology Europe(吉利技术欧洲公司) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing (LIESMARS)(信息工程测绘与遥感国家重点实验室(LIESMARS)) Wuhan University(武汉大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 提出J-LAW,通过耦合因子图联合优化度量物体位姿、潜在世界状态和潜在地标嵌入,实现定位与可操作世界建模的协同提升,实验证明能显著降低潜在预测误差和端点漂移。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22164 2026-09-01 cs.LG cs.RO 版本更新 90%

World Model Control by Trajectory Reachability Metrics

超越欧几里得距离:通过地平线匹配轨迹可达性度量修复潜在世界模型

Liangyu Li, Shengzhi Wang, Libin Qiu, Mingliang Xiong, Qingwen Liu

机构 * Tongji University(同济大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出轨迹可达性度量(TRM)作为固定潜在世界模型的后处理终端排名方法,通过训练小的成对头部来改进终端排名,从而提高连续操控任务的性能。

Comments 24 pages, including appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29904 2026-09-01 cs.CV 新提交 90%

Off-Manifold Refinement: Guiding Video Generators with a Frozen World Model

流形外细化:用冻结世界模型引导视频生成器

Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh

机构 * Fulbright University Vietnam(富布赖特越南大学) University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 该研究提出流形外细化(OMR)方法,用冻结的V-JEPA 2.1世界模型引导视频生成器,在低额外成本下提升了视频的物理一致性与语义 adherence指标。

Comments Accepted at BMVC 2026. Project page: https://itruonghai.github.io/omr

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30530 2026-09-01 cs.CL cs.SE 新提交 90%

WebWorld: The Browser as a World Model for Self-Improving Web Code

WebWorld:将浏览器作为自改进网页代码的世界模型

Jiajun Wu, Jian Yang, Yaxin Du, Wei Zhang, Haowen Wang, Junhang Cheng, Yuxuan Zhang, Tuney Zheng, Xianglong Liu, Ming Zhou

机构 * Beihang University(北京航空航天大学) Shanghai Jiao Tong University(上海交通大学) IQuest Research(IQuest研究院) Langboat

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 本研究提出WebWorld,将浏览器作为VLM无法欺骗的网页代码世界模型,通过浏览器签发合格证书的机制实现VLM自主交互与监督,使WebWorld-27B在两个基准任务上实现显著性能提升,达到强前沿系统水平。

Comments EMNLP Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25532 2026-09-01 gr-qc 版本更新 83%

A New Self-Dual Gravitational Instanton Solution on a Local Conformal Kählerian Manifold in a Brane World Model

一种在膜世界模型中局部共形凯勒流形上的新自对偶引力瞬子解

Reinoud Jan Slagter

专题命中 通用世界模型 :world model(title);world model(title)

AI总结 本文在共形膨胀子引力中找到了一种真空类克尔扭曲时空上的精确引力瞬子解,该解由一阶偏微分方程导出,与自对偶性相关,其奇点由五次多项式决定,拓扑为S^3×R/Z_2,并利用克莱因瓶上的对径边界条件描述了霍金粒子蒸发过程,最后揭示了与Janis-Newman-Winicour模型之间的联系。

Comments Version V4 : final. pictures improved --- spelling improved --- comments welcome---73 pages---70 pictures. 1 picture improved Aug 30

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03485 2026-09-01 cs.CV cs.AI cs.RO 版本更新 83%

Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion

Phys4D: 从视频扩散模型实现细粒度物理一致的4D建模

Haoran Lu, Shang Wu, Songling Liu, Jianshu Zhang, Maojiang Su, Guo Ye, Chenwei Xu, Lie Lu, Pranav Maneriker, Fan Du, Zhaoran Wang, Han Liu

机构 * Northwestern University(西北大学) Dolby Laboratories(杜比实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出Phys4D流水线,通过三阶段训练(伪监督预训练、物理监督微调、强化学习校正)从视频扩散模型学习物理一致的4D世界表示,显著提升细粒度时空与物理一致性。

Comments v2:Expanded the experiment section with more baselines and add more experiments in supplementary--corrected some typographical errors, and corrected author-affiliation information that was inaccurate in the previous version

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22642 2026-09-01 cs.LG cs.AI 版本更新 82%

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

Mol-JEPA:用于分子的多模态联合嵌入预测架构

Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff

机构 * University of Tübingen(蒂宾根大学) Boehringer Ingelheim(勃林格殷格翰) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Brown University(布朗大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对分子基础模型的化学无效增强等局限,提出Mol-JEPA多模态框架,利用模态掩码融入生化上下文,在基准测试中展现出优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02263 2026-09-01 cs.CV cs.AI 版本更新 82%

Social-JEPA: Emergent Geometric Isomorphism

Social-JEPA:涌现的几何同构

Haoran Zhang, Youjin Wang, Yi Duan, Rong Fu, Dianyu Zhao, Sicheng Fan, Shuaishuai Cao, Wentao Guo, Xiao Zhou

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 Social-JEPA通过让不同视角的独立代理学习环境模型,发现其潜在空间近似线性同构,从而实现跨代理的透明转换与高效迁移学习。

Comments Due to an unresolved dispute among the authors regarding the correctness of the experimental results and the validity of the conclusions, the team has decided to withdraw this paper until these issues can be fully resolved

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15281 2026-09-01 cs.CV 版本更新 69%

StableWorld: Towards Stable and Consistent Long Interactive Video Generation

StableWorld: 向稳定和一致的长交互视频生成迈进

Ying Yang, Zhengyao Lv, Yujia Zeng, Tianlin Pan, Haofan Wang, Yueming Lyu, Binxin Yang, Hubery Yin, Chen Li, Jing Lyu, Ziwei Liu, Chenyang Si

机构 * PRLab, NJU(南京大学PRLab) HKU(香港大学) UCAS(中国科学技术大学) WeChat, Tencent Inc.(腾讯公司) NTU(国立清华大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 StableWorld通过动态帧淘汰机制提升长交互视频生成的稳定性和时间一致性,适用于多种生成框架。

详情

展开后加载摘要…

URL PDF HTML 收藏