arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-09-01 至 2026-09-01 共收录 20 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 13 篇

2607.23602 2026-09-01 cs.RO cs.AI cs.LG 版本更新 94%

Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models

物理空间中相邻集的动作优于世界模型中的最佳预测

Liangyu Li, Qingwen Liu, Mingqing Liu, Wen Fang

机构 * Tongji University(同济大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究基于采样和潜在世界模型的控制器存在的条件失败提议过度生成问题,提出相邻集动作重建(ASAR)方法,在携带和释放评估集上,相比匹配选择提高了事件完成成功率,还刻画了选择风险。

Comments 23 pages, 7 figures. Includes supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03400 2026-09-01 cs.LG cs.AI 版本更新 94%

Better World Models Can Lead to Better Post-Training Performance

更好的世界模型可以导致更好的训练后性能

Prakhar Gupta, Henry Conklin, Sarah-Jane Leslie, Andrew Lee

机构 * University of Michigan(密歇根大学) Princeton University(普林斯顿大学) Harvard University(哈佛大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本研究通过比较不同世界建模策略,发现显式建模能提升Transformer的状态表示质量,从而增强强化学习后训练的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03603 2026-09-01 cs.CV cs.CL 版本更新 93%

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

世界模型遇见语言模型:论具体推理与抽象推理的互补性

Yucheng Zhou, Wei Tao, Yiwen Guo, Jianbing Shen

机构 * Nanyang Technological University(南洋理工大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出受控具体推理框架及PF-OPSD方法,通过结合世界模型的视觉模拟与多模态大语言模型的抽象推理,在空间前瞻和开放域物理预测任务上提升性能与鲁棒性。

Comments EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27367 2026-09-01 cs.CV cs.AI 版本更新 93%

Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

连续容量增长:面向JEPA世界模型的视觉Transformer编码器的任务复杂度驱动的宽度与深度扩展

Frederik Berenz

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文提出SCG方法,使JEPA世界模型的视觉Transformer编码器可按需连续扩展宽度或深度,在多任务上提升性能并实现更高参数效率,且零误扩展。

Comments 12 pages, 2 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15156 2026-09-01 cs.RO cs.AI 版本更新 93%

Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models

用于学习世界模型中反事实滚动的低秩动力学有效潜在载体

Yang Liu, Yuming Chen

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文针对学习世界模型的反事实滚动问题,提出基于低秩动力学有效潜在载体的方法,在两物体碰撞环境中验证秩4修补块可实现稳定的12步自主反事实滚动,且效果可复现。

Comments Revised manuscript with expanded Joint-intervention, carrier-relative, and recurrent-dynamics analyses. 54 pages, 7 figures. Code and data are available at https://github.com/lysea8282/dynamic-effective-latent-carriers

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12564 2026-09-01 cs.LG 版本更新 92%

Scaling Automatic Research Agents via World Models

通过世界模型扩展自动研究智能体

Xiyuan Yang, Sheikh Sarwar, Jingru Cheng, Zhan Shi, Duanshun Li, Huiyuan Chen, Haiyang Zhang, Xing Fan, Chenlei Guo, Jingrui He, Zhenyu Liao

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊公司)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 针对自动研究智能体扩展时的训练瓶颈,提出WMRL方法,结合两种缓解措施提升收敛性,训练加速3-4倍且性能优于更大规模智能体,还可迁移至多类任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28712 2026-09-01 cs.RO cs.LG 版本更新 91%

J-LAW: Joint Localization and Action-Conditioned World Modeling via Coupled Latent Factor Graphs

J-LAW:通过耦合潜在因子图实现联合定位与可操作世界建模

Guanqun Cao, Liang Chen

机构 * Geely Technology Europe(吉利技术欧洲公司) State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing (LIESMARS)(信息工程测绘与遥感国家重点实验室(LIESMARS)) Wuhan University(武汉大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 提出J-LAW,通过耦合因子图联合优化度量物体位姿、潜在世界状态和潜在地标嵌入,实现定位与可操作世界建模的协同提升,实验证明能显著降低潜在预测误差和端点漂移。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22164 2026-09-01 cs.LG cs.RO 版本更新 90%

World Model Control by Trajectory Reachability Metrics

超越欧几里得距离:通过地平线匹配轨迹可达性度量修复潜在世界模型

Liangyu Li, Shengzhi Wang, Libin Qiu, Mingliang Xiong, Qingwen Liu

机构 * Tongji University(同济大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出轨迹可达性度量(TRM)作为固定潜在世界模型的后处理终端排名方法,通过训练小的成对头部来改进终端排名,从而提高连续操控任务的性能。

Comments 24 pages, including appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25532 2026-09-01 gr-qc 版本更新 83%

A New Self-Dual Gravitational Instanton Solution on a Local Conformal Kählerian Manifold in a Brane World Model

一种在膜世界模型中局部共形凯勒流形上的新自对偶引力瞬子解

Reinoud Jan Slagter

专题命中 通用世界模型 :world model(title);world model(title)

AI总结 本文在共形膨胀子引力中找到了一种真空类克尔扭曲时空上的精确引力瞬子解,该解由一阶偏微分方程导出,与自对偶性相关,其奇点由五次多项式决定,拓扑为S^3×R/Z_2,并利用克莱因瓶上的对径边界条件描述了霍金粒子蒸发过程,最后揭示了与Janis-Newman-Winicour模型之间的联系。

Comments Version V4 : final. pictures improved --- spelling improved --- comments welcome---73 pages---70 pictures. 1 picture improved Aug 30

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03485 2026-09-01 cs.CV cs.AI cs.RO 版本更新 83%

Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion

Phys4D: 从视频扩散模型实现细粒度物理一致的4D建模

Haoran Lu, Shang Wu, Songling Liu, Jianshu Zhang, Maojiang Su, Guo Ye, Chenwei Xu, Lie Lu, Pranav Maneriker, Fan Du, Zhaoran Wang, Han Liu

机构 * Northwestern University(西北大学) Dolby Laboratories(杜比实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出Phys4D流水线,通过三阶段训练(伪监督预训练、物理监督微调、强化学习校正)从视频扩散模型学习物理一致的4D世界表示,显著提升细粒度时空与物理一致性。

Comments v2:Expanded the experiment section with more baselines and add more experiments in supplementary--corrected some typographical errors, and corrected author-affiliation information that was inaccurate in the previous version

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22642 2026-09-01 cs.LG cs.AI 版本更新 82%

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

Mol-JEPA:用于分子的多模态联合嵌入预测架构

Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff

机构 * University of Tübingen(蒂宾根大学) Boehringer Ingelheim(勃林格殷格翰) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Brown University(布朗大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对分子基础模型的化学无效增强等局限,提出Mol-JEPA多模态框架,利用模态掩码融入生化上下文,在基准测试中展现出优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02263 2026-09-01 cs.CV cs.AI 版本更新 82%

Social-JEPA: Emergent Geometric Isomorphism

Social-JEPA:涌现的几何同构

Haoran Zhang, Youjin Wang, Yi Duan, Rong Fu, Dianyu Zhao, Sicheng Fan, Shuaishuai Cao, Wentao Guo, Xiao Zhou

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 Social-JEPA通过让不同视角的独立代理学习环境模型,发现其潜在空间近似线性同构,从而实现跨代理的透明转换与高效迁移学习。

Comments Due to an unresolved dispute among the authors regarding the correctness of the experimental results and the validity of the conclusions, the team has decided to withdraw this paper until these issues can be fully resolved

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15281 2026-09-01 cs.CV 版本更新 69%

StableWorld: Towards Stable and Consistent Long Interactive Video Generation

StableWorld: 向稳定和一致的长交互视频生成迈进

Ying Yang, Zhengyao Lv, Yujia Zeng, Tianlin Pan, Haofan Wang, Yueming Lyu, Binxin Yang, Hubery Yin, Chen Li, Jing Lyu, Ziwei Liu, Chenyang Si

机构 * PRLab, NJU(南京大学PRLab) HKU(香港大学) UCAS(中国科学技术大学) WeChat, Tencent Inc.(腾讯公司) NTU(国立清华大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 StableWorld通过动态帧淘汰机制提升长交互视频生成的稳定性和时间一致性,适用于多种生成框架。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频世界模型 1 篇

2606.01600 2026-09-01 cs.CV cs.CL cs.RO 版本更新 96%

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

RoboTrustBench:机器人操作视频世界模型的可信度基准测试

Huiqiong Li, Jiayu Wang, Zhiting Mei, Anirudha Majumdar, Jingjing Chen, Bin Zhu

机构 * Singapore Management University(新加坡国立管理学院) Fudan University(复旦大学) Princeton University(普林斯顿大学)

专题命中 视频世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 针对视频世界模型在机器人操作中的可信度问题,提出RoboTrustBench基准,包含正常、约束敏感、反事实和对抗四种场景,通过专家验证的指令-图像对和六维评估协议,发现当前模型在约束推理、反事实基础、物理交互和不安全指令抑制方面存在不足。

Comments EMNLP 2026 Findings, Project: https://huiqiongli.github.io/RoboTrustBench/

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 具身与机器人 3 篇

2608.24101 2026-09-01 cs.RO 版本更新 86%

TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks

TrAct:通过视觉轨迹连接机器人控制与视觉预测

Zhi Cao, Howard Ji, Kevin Zhang, Kuangzhi Ge, Li Fei-Fei, Jiajun Wu, Huang Huang

机构 * University of Michigan(密歇根大学) Stanford University(斯坦福大学)

专题命中 具身与机器人 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 TrAct是基于世界模型的机器人决策框架,以视觉轨迹为控制与预测的中间接口,在LIBERO-INTEGRAL基准和Franka任务上,相较基线π₀.5显著提升了操作成功率与视频预测质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22067 2026-09-01 cs.RO cs.AI cs.CV cs.LG 版本更新 75%

DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation

从未来隐状态推断机器人操纵动作

Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Peilin Huang, Han Fu, Zhuo Li, Xiaoxue Ren

机构 * DeepLeap Research(DeepLeap研究院)

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG、cs.CV

AI总结 本文提出无需视频生成的机器人操纵模型DELE-w0.5,通过从捕获动作相关物理结果的未来隐状态推断动作,在4项长程操纵任务的480次试验中,其性能优于最强基线47.5和30.7个百分点,实现最优表现。

Comments DeepLeap Technology Co., Ltd., Shenzhen, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13073 2026-09-01 cs.RO cs.CV 版本更新 71%

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

用于机器人操作的基于大型VLM的视觉-语言-动作模型:综述

Rui Shao, Wei Li, Lingsen Zhang, Renshan Zhang, Zhiyang Liu, Ran Chen, Liqiang Nie

机构 * School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(计算机科学与技术学院,哈尔滨工业大学(深圳))

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.CV、cs.RO

AI总结 本综述首次系统分类梳理用于机器人操作的基于大型VLM的VLA模型,明确其定义与两类架构,考察其与先进领域的集成等内容,整合进展并提供更新项目页面

Comments Under Minor Revision at IEEE TPAMI, Project Page: https://github.com/JiuTian-VL/Large-VLM-based-VLA-for-Robotic-Manipulation

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 模型式强化学习 2 篇

2605.00272 2026-09-01 q-bio.QM 版本更新 65%

LNODE: latent dynamics reveal the shared spatiotemporal structure of amyloid-$β$ progression

LNODE:潜变量揭示阿尔茨海默病β淀粉样蛋白进展的共享时空结构

Zheyu Wen, George Biros

专题命中 模型式强化学习 :latent dynamics(title)

AI总结 LNODE模型通过PET影像校准,揭示阿尔茨海默病β淀粉样蛋白进展的时空结构,具备融合、定量分析和解释能力,展现高参数可识别性和稳定性。

Comments 36 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23188 2026-09-01 cs.LG physics.flu-dyn 版本更新 50%

Efficient Real-Time Adaptation of ROMs for Unsteady Flows Using Data Assimilation

利用数据同化高效实时适应不稳流的ROM

Ismaël Zighed, Andrea Nóvoa, Luca Magri, Taraneh Sayadi

专题命中 模型式强化学习 :latent dynamics(abstract);分类 cs.LG

AI总结 本文提出一种基于数据同化的高效ROM重训练方法,通过概率性VAE和变压器网络,在稀疏观测下实现快速实时适应不稳流的降阶模型。

Journal ref Computers & Fluids, Volume 319, 2026, 107250,

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 仿真与规划 1 篇

2601.04035 2026-09-01 cs.AI 版本更新 90%

MobileDreamer: Generative Sketch World Model for GUI Agent

MobileDreamer: 用于GUI代理的生成式草图世界模型

Yilin Cao, Yufeng Zhong, Zhixiong Zeng, Siran Dai, Liming Zheng, Jing Huang, Haibo Qiu, Peng Shi, Wenji Mao

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Meituan(美团)

专题命中 仿真与规划 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 MobileDreamer通过生成式草图世界模型和rollout想象策略,提升GUI代理在长周期任务中的决策能力,任务成功率提升5.25%。

详情

展开后加载摘要…

URL PDF HTML 收藏