arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-05-12 至 2026-05-12 共收录 20 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 20 篇

2605.10858 2026-05-12 cs.CV cs.RO 95%

Is Your Driving World Model an All-Around Player?

你的驾驶世界模型是全能选手吗?

Lingdong Kong, Ao Liang, Tianyi Yan, Hongsi Liu, Wesley Yang, Ziqi Huang, Xian Sun, Wei Yin, Jialong Zuo, Yixuan Hu, Dekai Zhu, Dongyue Lu, Youquan Liu, Guangfeng Jiang, Linfeng Li, Xiangtai Li, Long Zhuo, Lai Xing Ng, Benoit R. Cottereau, Changxin Gao, Liang Pan, Wei Tsang Ooi, Ziwei Liu

专题命中 通用世界模型 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 本文提出WorldLens基准测试,评估驾驶世界模型在视觉和行为真实性方面的综合表现,揭示现有模型在不同维度上的不足,并引入WorldLens-26K和WorldLens-Agent提升评估的可解释性。

Comments CVPR 2026 VideoWorldModel Workshop; Project Page at https://worldbench.github.io/worldlens GitHub at https://github.com/worldbench/WorldLens

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08279 2026-05-12 cs.LG cs.AI 95%

LaWM: Least Action World Models for Long-Horizon Physical Consistency from Visual Observations

LaWM:基于视觉观测的最短作用世界模型用于长时间物理一致性

Qixin Xiao, Maani Ghaffari

机构 * University of Michigan(密歇根大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 LaWM通过在视觉潜在空间中实现最小作用原理,提升长horizon视觉预测的物理一致性,改进了物理不变性、背景一致性、运动平滑度和外观几何预测指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10806 2026-05-12 cs.CV cs.AI cs.LG 94%

PhyGround: Benchmarking Physical Reasoning in Generative World Models

PhyGround:用于生成世界模型中物理推理的基准测试

Juyi Lin, Arash Akbari, Yumei He, Lin Zhao, Haichao Zhang, Arman Akbari, Xingchen Xu, Zoe Y. Lu, Enfu Nan, Hokin Deng, Edmund Yeh, Sarah Ostadabbas, Yun Fu, Jennifer Dy, Pu Zhao, Yanzhi Wang

机构 * Northeastern University(东北大学) Tulane University(路易斯安那州立大学) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 PhyGround通过250个精心设计的提示和13种物理定律的分类,评估视频生成中的物理推理能力,采用大规模人工研究验证并发布专用VLM评估工具PhyJudge-9B,显著降低偏见。

Comments Preprint. 56 pages, 39 figures, 40 tables. Project page: https://phyground.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09241 2026-05-12 cs.LG cs.AI 94%

Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models

Sub-JEPA:子空间高斯正则化用于稳定的端到端世界模型

Kai Zhao, Dongliang Nie, Yuchen Lin, Zhehan Luo, Yixiao Gu, Deng-Ping Fan, Dan Zeng

机构 * Shanghai University(上海大学) The University of Manchester(曼彻斯特大学) Nankai University(南开大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 Sub-JEPA通过在多个随机子空间中应用高斯约束,缓解了JEPA训练中的偏差-方差权衡问题,提升了训练稳定性与表示灵活性,实验显示其在连续控制环境中优于LeWM。

Comments https://github.com/intcomp/Sub-JEPA

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08578 2026-05-12 cs.LG cs.AI 94%

Probing the Impact of Scale on Data-Efficient, Generalist Transformer World Models for Atari

探测数据高效、通用型Transformer世界模型在Atari上的尺度影响

Jooyeon Kim

机构 * Graduate School of Artificial Intelligence(人工智能研究生院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文通过最小化Transformer世界模型分析Atari 100k基准的缩放行为,发现环境存在不同的缩放区域,联合训练可稳定缩放动态并提升下游控制性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08954 2026-05-12 cs.LG cs.AI 93%

MolWorld: Molecule World Models for Actionable Molecular Optimization

MolWorld: 用于可操作分子优化的分子世界模型

Yang Qiao, Bo Pan, Hao-Wei Pang, Peter Zhiping Zhang, Liying Zhang, Liang Zhao

机构 * Emory University(埃默里大学) Merck & Co., Inc.(默克公司)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 MolWorld通过分子转移图的序列扩展实现可操作分子优化,通过学习世界模型增强分子结构连接性,提升分子设计的可行性和连续性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09886 2026-05-12 cs.RO 92%

Network-Efficient World Model Token Streaming

网络高效的世界模型令牌流

Shatadal Mishra, Ahmadreza Moradipari, Nejib Ammar

机构 * InfoTech Labs, Toyota Motor North America R\&D, Mountain View, CA, USA

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);driving world model(abstract)

AI总结 本文研究了在分布式计算和连接车辆中高效传输离散世界模型状态的方法,提出了一种自适应算法,通过余弦距离优先更新令牌并触发关键帧,从而在带宽受限环境下提升同步效率和下游令牌动态的实用性。

Comments Accepted at IEEE VNC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11537 2026-05-12 cs.LG cs.AI 90%

Simulus: Combining Improvements in Sample-Efficient World Model Agents

Simulus:结合高效样本世界模型代理的改进

Lior Cohen, Kaixin Wang, Bingyi Kang, Uri Gadot, Shie Mannor

机构 * Technion(技术ion大学) Microsoft Research(微软研究院) ByteDance Seed(字节跳动种子实验室)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出Simulus,一种模块化的基于标记的世界模型代理,通过整合灵活的标记化框架、内在动机、优先世界模型回放和回归-分类方法,提升了样本效率,适用于多个基准测试。

Comments Revised version: updated title, abstract, and framing to better reflect our contributions and situate the work within the literature

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09131 2026-05-12 cs.AI cs.MA 90%

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

MCP-Cosmos:用于MCP环境复杂任务执行的世界模型增强型智能体

Giridhar Ganapavarapu, Dhaval Patel

机构 * IBM

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 MCP-Cosmos融合世界模型与MCP框架,通过引入生成世界模型提升智能体在复杂任务中的执行能力,实验显示其在工具成功率和参数准确性等方面有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08268 2026-05-12 cs.MA cs.AI 87%

Insider Attacks in Multi-Agent LLM Consensus Systems

多智能体大语言模型共识系统中的内部攻击

Xiaolin Sun, Zixuan Liu, Yibin Hu, Zizhan Zheng

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,Location,Country) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,Location,Country) Department of Computer Science, Tulane University, New Orleans, United States of America(计算机科学系, Tulane大学,新奥尔良,美国)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 研究多智能体大语言模型共识系统中的内部攻击问题,提出基于世界模型的框架,通过学习良性智能体的潜在行为状态并利用强化学习训练攻击者,有效降低共识率并延长分歧时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10404 2026-05-12 cs.CV 81%

Position: Life-Logging Video Streams Make the Privacy-Utility Trade-off Inevitable

位置:生命记录视频流使隐私与效用的权衡不可避免

Tianyuan Zou, Liang Yue, Yang Liu, Ya-Qin Zhang, Sijie Cheng

机构 * Institute for AI Industry Research, Tsinghua University, Beijing, China(清华大学人工智能产业研究院) RayNeo.AI, Shenzhen, China(深圳RayNeo.AI) Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 生命记录视频流虽提升AI系统效用,但暴露敏感信息引发隐私风险,现有保护措施无法兼顾效用与隐私,需进一步研究权衡设计与标准化评估。

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08956 2026-05-12 cs.AI 81%

Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

代理式AI科学家并非为自主科学发现而建

Harshit Bisht, Vinay Kumar, Kevin Maik Jablonka, Mausam, N. M. Anoop Krishnan

机构 * Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(印度理工学院德里人工智能学院) Department of Computer Science and Engineering, Indian Institute of Technology Delhi(印度理工学院德里计算机科学与工程系) Department of Civil and Environmental Engineering, Indian Institute of Technology Delhi(印度理工学院德里土木与环境工程系) Laboratory of Organic and Macromolecular Chemistry (IOMC), Friedrich Schiller University Jena(耶拿弗里德里希·席勒大学有机与大分子化学实验室) Center for Energy and Environmental Chemistry Jena, Friedrich Schiller University Jena(耶拿弗里德里希·席勒大学能源与环境化学中心) Helmholtz Institute for Polymers in Energy Applications Jena (HIPOLE Jena)(耶拿海德堡聚合物能源应用研究所(HIPOLE耶拿))

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文指出,尽管代理式AI科学家能作为合作者,但其设计并非为自主科学发现。挑战包括问题选择偏差、实验室知识缺失、输出多样性压缩及缺乏实验反馈,需重新审视基础设计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23629 2026-05-12 cs.GR 80%

From Visual Synthesis to Interactive Worlds: Toward Production-Ready 3D Asset Generation

从视觉合成到交互世界:迈向可生产3D资产生成

Jiafeng Wu, Zhuofan Lou, Jian Liu, Dazhao Du, Chunchao Guo, Song Guo

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文探讨了从孤立视觉形状到可部署的结构化资产的3D内容生成进展,分析了游戏开发、AI、世界模拟等对3D资产的需求,提出基于资产生产流程的分类方法,评估现有方法的生成质量和可用性,并指出数据质量、生成可控性等挑战。

Comments Preprint. Jiafeng Wu and Zhuofan Lou contributed equally. Project page: https://christinebobby.github.io/production-ready-3d-survey/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09693 2026-05-12 cs.CV cs.AI cs.LG 73%

Do multimodal models imagine electric sheep?

多模态模型是否想象出电羊?

Santhosh Kumar Ramakrishnan, Carl Vondrick, Raja Giryes, Philipp Krähenbühl, Vladlen Koltun

机构 * Apple(苹果公司)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG、cs.CV

AI总结 研究发现多模态模型在解决空间谜题时会形成心理图像,通过微调Qwen3.5 VLM解决多种视觉推理任务,发现动作序列预测能提升解谜准确率,尤其在需要空间推理的任务中效果显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09790 2026-05-12 cs.DC cs.AI cs.LG 71%

Multi-Tier Labeling and Physics-Informed Learning for Orbital Anomaly Detection at Scale

多层级标签与物理引导学习用于大规模轨道异常检测

Yong Fu

机构 * Substratum Labs, Inc(Substring实验室)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出多层级标签方法,结合物理规则、IMM-UKF和补充元素校准,实现大规模轨道异常检测,通过Transformer模型提升召回率,探讨神经ODE在轨道世界模型中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09650 2026-05-12 cs.AI cs.LG 71%

Workspace Optimization: How to Train Your Agent

工作区优化:如何训练你的智能体

Elad Sarafian, Gal Kaplun, Ron Banner, Daniel Soudry, Boris Ginsburg

机构 * NVIDIA

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过优化智能体的工作区结构来提升其在多轮任务中的表现,通过模拟权重空间训练方法,利用人工制品、证据、反例和文本反馈来改进智能体的执行能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09678 2026-05-12 cs.AI 69%

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities

荒诞世界:一种简单却强大的方法,用于将现实世界扭曲以探测LLM推理能力

Ryan Albright, Golam Md Muktadir, Zarif Ikram, S M Jubaer, Mehrab Hossain, Dianbo Liu

机构 * The Nueva School(新维学校) University of Southern California(南加州大学) Notre Dame College(诺特大学) Arizona State University(亚利桑那州立大学) National University of Singapore(新加坡国立大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 本文提出Absurd World框架,通过扭曲现实世界来测试LLM的推理能力,验证其在简单逻辑任务中的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09355 2026-05-12 cs.LG 69%

FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

FLAME:适应性专家混合用于连续多模态多任务学习

Xing Han, Shravan Chaudhari, Tanvi Ranade, Rama Chellappa, Suchi Saria

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 FLAME提出了一种可扩展的专家混合框架,支持多任务预训练和连续学习,通过模态特定路由器处理不同模态的任务,同时利用低秩内存子空间压缩专家知识,提升参数效率和减少灾难性遗忘。

Comments 37 pages, 25 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.04899 2026-05-12 cs.LG 69%

A geometric relation of the error introduced by sampling a language model's output distribution to its internal state

语言模型输出分布采样引入的误差与内部状态的几何关系

Albert F. Modenbach

机构 * Department of Mathematics, King's College London, United Kingdom(伦敦国王学院数学系)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 研究揭示语言模型输出分布采样误差与内部状态几何结构的关系,通过几何方法分析token嵌入空间,发现其曲率在棋类任务中反映模型对问题的内部表示。

Comments 12 Pages, 10 Figures, 2 Appendices. To appear in Proceedings of ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08958 2026-05-12 cs.LG 60%

Learning predictive models for combinations of heterogeneous proteomic data sources

学习异质蛋白质组数据源的预测模型

Michal Valko, Richard Pelikan, Miloš Hauskrecht

专题命中 通用世界模型 :predictive model(title);predictive models(title);分类 cs.LG

AI总结 本文研究了两种蛋白质混合物表达数据源的个体和联合效用,提出了一种模型融合方法以最大化异质数据集的结合效益。

Comments Published at in AMIA Summit on Translational Bioinformatics (STB 2008

详情

展开后加载摘要…

URL PDF HTML 收藏