arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-08-26 至 2026-08-26 共收录 25 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 14 篇

2608.24855 2026-08-26 cs.CV 新提交 94%

LeFlow: Generative Latent Flow Planning for World Models

LeFlow:面向世界模型的生成式潜在流规划

Hsiang-Wei Huang, Jianxu Shangguan, Junbin Lu, Jenq-Neng Hwang

机构 * University of Washington(华盛顿大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出LeFlow,它从世界模型学习可复用潜在轨迹先验,将规划转为条件潜在轨迹生成,在四个目标条件像素控制基准中提升规划成功率并大幅减少规划时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24885 2026-08-26 cs.RO cs.CV 新提交 94%

Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning

机器人世界模型真的会遵循动作吗?面向策略学习的动作条件生成诊断与对齐

Sixiang Chen, Jiaming Liu, Jixian Wu, Yichen Guo, Tinghao Wang, Siyuan Qian, Hao Chen, Jiajun Cao, Jian Tang, Shanghang Zhang

机构 * Peking University(北京大学) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) New York University(纽约大学) University of Electronic Science and Technology of China(电子科技大学) Nanyang Technological University(南洋理工大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文针对动作条件世界模型的动作遵循问题,提出WorldEcho诊断工具与WorldSync改进方法,经实验验证WorldSync可提升动作遵循性能,作为可靠模拟器助力策略改进并提高任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23720 2026-08-26 cs.CV 新提交 93%

Platonic Representation Hypothesis on World Models

世界模型的柏拉图式表征假说

Wenhow Li (1), Chengwei MA (1), Hui Xiong (1), Ying-Cong Chen (1), Lei Zhang (1) ((1) The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China)

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究针对世界模型的表征性质,提出预测一致性假设,通过DINO-WM实验发现性能良好的世界模型会形成几何相似的内部结构,且模型特征可跨模型映射,证实预测一致性能促进共享潜在结构的形成。

Comments 18 pages, 10 figures, 2 tables. Wenhow Li and Chengwei MA contributed equally. Project page: this https URL (https://sellerbubble.github.io/platonic-representation-hypothesis-on-world-models/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09696 2026-08-26 cs.AI 版本更新 93%

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

模型发现智能体:用于数据高效发现机制世界模型的大语言模型辅助贝叶斯实验设计

Kevin Murphy

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究提出模型发现智能体(MDA),结合LLM与贝叶斯机制,在少量干预下发现机制世界模型,在三类基准上实现数据高效模型学习与可靠干预预测的SOTA性能。

Comments v4: Major update! Fixed a leak in the prompts for physics and chemistry benchmarks and re-ran experiments (fortunately results did not change much), added Boxing Gym benchmark (requires generating NumPyro code), significantly simplified the figures and evaluation protocol, reframed the narrative around SMC^3, polished the presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24680 2026-08-26 cs.CV 新提交 93%

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training

Game2World引擎:解锁野外游戏视频用于世界模型训练

Wenxuan Shen, Dongna Jin, Dongping Chen

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world-model(abstract)

AI总结 针对游戏视频界面干扰世界模型训练的问题,提出GameUI-Taxonomy与G2WEngine框架构建Game2World数据集,研发无掩码UI去除模型GameCleaner,提升了世界模型训练效果与UI去除性能。

Comments We are currently building Gaming World Model and data engine that transfers game dynamics to robotics. Feel free to contact Dongping Chen (dongpingchen0612@gmail.com) if you are interested in research collaboration or financial support

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07265 2026-08-26 math.OC cs.LG 交叉投稿 91%

Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control

用于控制的学习型世界模型中的度量非崩溃:逼近理论、有限样本几何保证与确定性规划迁移

Alain Bensoussan, Minh-Nhat Phung, Minh-Binh Tran

专题命中 通用世界模型 :world model(title);world models(title);world model(title);world models(title)

AI总结 该研究为非线性确定性控制系统的学习型世界模型构建三部分数学理论,含逼近理论、有限样本几何保证及确定性规划迁移方法,通过数值实验验证了相关方法的有效性。

Comments Revised version prepared in response to the editorial assessment. The main manuscript is 32 pages; detailed mathematical derivations have been moved to the accompanying Supplementary Material. The principal results and contributions are strengthened and clarified

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21114 2026-08-26 cs.CV cs.AI 版本更新 88%

CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents

CIVA:面向视觉世界模型智能体的评论者诱导价值子空间攻击

Jiancheng Wang, Mingli Zhu, Tong Zhang, Jiaqi Ruan, Wei Wang, Siyuan Liang, Dacheng Tao

专题命中 通用世界模型 :world-model(title,abstract);world-model(title,abstract);分类 cs.AI、cs.CV

AI总结 该研究针对视觉世界模型智能体提出CIVA攻击方法,通过提取价值子空间优化扰动,在多个基准任务上优于现有方法,实现了低时间变化下的显著奖励下降。

Comments Includes supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23790 2026-08-26 cs.CV q-bio.NC 新提交 81%

Primate vision reveals a missing principle for robust dynamic AI

灵长类视觉揭示了鲁棒动态人工智能缺失的原理

Matteo Dunnhofer, Christian Micheloni, Kohitij Kar

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 该研究对比人类与猕猴视觉及相关神经网络,发现预测世界模型兼具跨外观泛化与高神经保真度,提出将运动逐步整合到物体表征是鲁棒动态视觉的关键原理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13554 2026-08-26 cs.ET cs.IT cs.RO 版本更新 81%

Observability Engineering: From Measurement to Information Generation in Active Sensing Systems

从快照感知到持久电磁世界建模:一种面向ISAC的生成空间视角

Pin-Han Ho, Limei Peng

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出一种基于生成空间的毫米波感知框架,通过低维激励空间实现灵活、可扩展的持续电磁世界建模,避免了快照感知的局限性。

Comments 7 pages, 6 figures/tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12144 2026-08-26 cs.CV cs.RO eess.IV 版本更新 71%

O3N: Omnidirectional Open-Vocabulary Occupancy Prediction for Embodied Intelligent Robotics

O3N: 全向开放词汇占用预测

Mengfei Duan, Hao Shi, Fei Teng, Guoqiang Zhao, Yuheng Zhang, Zhiyong Li, Kailun Yang

机构 * Hunan University(湖南大学) Zhejiang University(浙江大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV、cs.RO

AI总结 O3N通过全向感知和开放词汇方法,实现三维空间的连续表示和长距离上下文建模,提升具身智能在开放世界中的感知与建模能力。

Comments The source code will be made publicly available at this https URL (https://github.com/MengfeiD/O3N)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24534 2026-08-26 cs.AI 新提交 69%

Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models

面向生理安全临床语言模型的神经符号对齐

Abdulhady Abas Abdullah, Erik Cambria, Milena Zivkovic

机构 * University of Kurdistan Hewlêr(库尔德斯坦大学) Nanyang Technological University(南洋理工大学) Massachusetts Institute of Technology(麻省理工学院) University of Kragujevac(克拉古耶瓦茨大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 提出神经符号对齐框架,耦合临床LLM与HGNN生理世界模型,在CSB基准上显著提升临床LLM的生理安全性能,优于ORPO、GPT-4等方法,为临床LLM安全对齐提供新路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23706 2026-08-26 cs.AI 新提交 69%

Do LLMs Understand Limit Order Book Dynamics?

大型语言模型(LLM)是否理解限价订单簿(LOB)动态?

Junxiao Chen, Paul Glasserman

机构 * Columbia Business School(哥伦比亚商学院)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 本研究测试基于LOB数据训练的LLM,发现其隐式世界模型未掌握LOB状态,导致预测有偏且存在虚假可预测性,为此将相关测试扩展至LOB的随机动态环境。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23383 2026-08-26 cs.CV 版本更新 69%

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

面向持续故事与交互世界的长时序视听生成

Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang

机构 * Joy Future Academy, JD(京东探索研究院)

专题命中 通用世界模型 :world-model(abstract);world-model(abstract);分类 cs.CV

AI总结 研究针对长时序视听生成的需求,提出JoyAI-Echo-1.5系统,含长视频与世界模型变体,通过专用技术实现跨镜头一致性等性能,在相关基准上取得领先结果,为生成连贯内容提供基础。

Comments Project page: this https URL (https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24056 2026-08-26 cs.LG cs.CE 新提交 67%

PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation

PhysicsBench:面向工程设计与仿真的生成式及预测式模型统一排行榜

Sang Won Lee, Hyogu Jeong, Namwoo Kang

专题命中 通用世界模型 :predictive model(title,abstract);predictive models(title,abstract);分类 cs.LG

AI总结 PhysicsBench是面向工程设计与仿真的统一基准排行榜,涵盖多维度任务与数据集,采用标准化流程评估66个模型,可实现模型去偏排名,助力模型选择。

Comments 40 pages, 12 figures, 8 tables. Leaderboard: this https URL (https://leaderboard.narnia.ai) | Data: this https URL (https://github.com/Narnialabs/leaderboard)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身与机器人 5 篇

2608.23863 2026-08-26 cs.RO 新提交 88%

DreamLedger: Execution-Settled Credit Files for World-Model Imagination in Robot Decision Loops

DreamLedger:机器人决策回路中用于世界模型想象的执行结算信用文件

Xianyao Li, Ruitong Tian, Rui Min, Fang Xu, Jing Du

机构 * University of Florida(佛罗里达大学)

专题命中 具身与机器人 :world-model(title,abstract);world-model(title,abstract);分类 cs.RO

AI总结 该研究提出DreamLedger执行结算信用文件,管控机器人世界模型预测的消耗,在多模拟域和真实机械臂上验证,可减少无效想象、降低验证探测次数,适用于多模型与硬件场景。

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01381 2026-08-26 cs.RO 版本更新 88%

DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation

DreamTrajectory:面向移动操作的、结合世界模型对齐的轨迹引导动作生成方法

Zheng Yang, Wenjie Zhang, Xiangyu Chen, Wenxuan Song, Xianpeng Wang, Yihang Kang, Jiawen Wen, Wen Chen, Lujia Wang, Renjing Xu, Haoang Li, Xiaowen Chu

专题命中 具身与机器人 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 DreamTrajectory是面向移动操作的轨迹引导框架,通过联合预测末端执行器轨迹与全身动作块、结合轨迹世界模型的测试时细化,在MS-HAB及真实任务中大幅提升操作成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24199 2026-08-26 cs.RO 新提交 86%

NVIDIA Cosmos-H-Dreams: Real-Time Generative Physics Simulation for Surgical Robotics

NVIDIA Cosmos-H-Dreams:面向外科机器人的实时生成式物理仿真

Javier Gamazo Tejero, Lukas Zbinden, Keyur Sheth, Raghavendra K M, Nadim Daher, Diego Granero Maraña, Filip Binkiewicz, Patrick Thornycroft, Mahdi Azizian, Sean D. Huver

专题命中 具身与机器人 :world model(abstract);world-model(abstract);video world model(abstract);world model(abstract)

AI总结 该研究提出Cosmos-H-Dreams系统,结合生成模型与蒸馏技术实现实时外科机器人仿真,支持多控制器驱动,为外科教育等提供开放基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24101 2026-08-26 cs.RO 新提交 86%

TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks

TrAct:通过视觉轨迹连接机器人控制与视觉预测

Zhi Cao, Howard Ji, Kevin Zhang, Kuangzhi Ge, Li Fei-Fei, Jiajun Wu, Huang Huang

机构 * University of Michigan(密歇根大学) Stanford University(斯坦福大学)

专题命中 具身与机器人 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 TrAct是基于世界模型的机器人决策框架,以视觉轨迹为控制与预测的中间接口,在LIBERO-INTEGRAL基准和Franka任务上,相较基线π₀.5显著提升了操作成功率与视频预测质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22067 2026-08-26 cs.RO cs.AI cs.CV cs.LG 版本更新 75%

Inferring Action from Future Latent State for Robotic Manipulation

从未来隐状态推断机器人操纵动作

Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Peilin Huang, Han Fu, Zhuo Li, Xiaoxue Ren

机构 * DeepLeap Research(DeepLeap研究院)

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG、cs.CV

AI总结 本文提出无需视频生成的机器人操纵模型DELE-w0.5,通过从捕获动作相关物理结果的未来隐状态推断动作,在4项长程操纵任务的480次试验中,其性能优于最强基线47.5和30.7个百分点,实现最优表现。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 自动驾驶 3 篇

2605.10426 2026-08-26 cs.CV cs.AI 版本更新 85%

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving

CoWorld-VLA:面向自动驾驶的多专家世界模型中的思考

Minqing Huang, Yujiao Xiang, Zihan Liang, Jiajie Huang, Jingqi Wang, Yuheng Zhou, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang, Gong Che

机构 * Afari Intelligent Drive(Afari智能驾驶公司) University of Electronic Science and Technology of China(电子科技大学) Shanghai Jiao Tong University(上海交通大学) Beijing University Of Posts and Telecommunications(北京邮电大学) Tianjin University(天津大学)

专题命中 自动驾驶 :world model(title);world model(title);分类 cs.AI、cs.CV

AI总结 本文提出CoWorld-VLA,通过多专家世界推理框架,利用显式条件指导动作规划,提升自动驾驶的场景生成与路径规划性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23486 2026-08-26 cs.CV cs.RO 版本更新 71%

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

GeoWAM:面向自动驾驶的视觉几何世界动作模型

Yiren Lu, Xin Ye, Jiaming Liu, Philip Jacobson, Jin Yao, Yi-chung Chen, Liam Merino, Dhruva Dixith Kurra, Min Cai, Tom Lampo, Yu Yin, Danhua Guo, Burhan Yaman

机构 * Uber AV Labs(优步自动驾驶实验室) Case Western Reserve University(凯斯西储大学)

专题命中 自动驾驶 :world model(abstract);world model(abstract);分类 cs.CV、cs.RO

AI总结 该研究提出 GeoWAM,一种用于自动驾驶的视觉几何世界动作模型,通过预训练预测未来场景几何来学习动力学,经评估其生成的驾驶策略比图像基方案更强,确立未来几何预测为自动驾驶有效预训练目标。

Comments Project page: this https URL (https://yiren-lu.com/project_pages/geowam/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24136 2026-08-26 cs.LG 新提交 50%

Steering Recurrent Reasoners at Inference Time with Readout Feedback

在推理阶段通过读出反馈引导循环推理器

Shunsuke Kamiya, Masanori Koyama, Seongcheol Jeong, Fumiya Uchiyama, Kenji Kubo, Kohei Hayashi, Masahiro Suzuki, Yutaka Matsuo

机构 * Graduate School of Engineering, The University of Tokyo(东京大学大学院工学系研究科)

专题命中 自动驾驶 :latent dynamics(abstract);分类 cs.LG

AI总结 该研究提出测试时干预方法RoFB,将中间预测转为耦合力注入循环模型隐动态,在数独、迷宫任务的三类循环模型上提升性能,且计算成本相当或更低。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 模型式强化学习 1 篇

2608.24044 2026-08-26 cs.LG 新提交 88%

XP-JEPA: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

XP-JEPA:用于可预测潜在动力学的交叉预测物理基础

Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi

专题命中 模型式强化学习 :latent dynamics(title,abstract);world model(abstract);world models(abstract);world model(abstract)

AI总结 XP-JEPA通过将视觉与物理表征交叉预测,提升了潜在动力学的可预测性,在多任务套件上降低了展开漂移并提高了控制成功率,无需特权输入即可实现更优的基于展开的控制性能。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 仿真与规划 2 篇

2603.14603 2026-08-26 cs.RO 版本更新 68%

Latent Dynamics-Aware OOD Monitoring for Trajectory Prediction with Provable Guarantees

隐式动态感知的领域外监测用于轨迹预测的可证明保障

Tongfei Guo, Lili Su

专题命中 仿真与规划 :latent dynamics(title);分类 cs.RO

AI总结 本文提出基于快速突变点检测的轨迹预测领域外监测方法,通过隐马尔可夫模型建模预测误差演化,实现无需显式知识的领域外检测并保证延迟和误报率的可证明保障。

Comments Accepted by 2026 IEEE International Conference on Automation Science and Engineering (CASE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23622 2026-08-26 cs.AI cs.CL cs.MA cs.SE 新提交 58%

LLM Agents Perform Controlled Experiments Using Simulation Models

大语言模型智能体使用模拟模型开展对照实验

Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes Stümpfle, Johannes Sigel, Akshay Narla, Gavin K. Reynolds, Anna Jawor-Baczynska, Pol Llopart

机构 * Institute for Industrial Automation and Software Engineering(工业自动化与软件工程研究所) University of Stuttgart(斯图加特大学) AstraZeneca(阿斯利康)

专题命中 仿真与规划 :simulation model(title,abstract);分类 cs.AI、cs.MA

AI总结 本研究提出多智能体框架,将LLM与高保真模拟模型结合,使LLM智能体可开展制药工艺设计的对照实验,生成更具体可操作的优化建议,在工业场景中表现更优。

Comments Accepted at the 31st IEEE International Conference on Emerging Technologies and Factory Automation ETFA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏