arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-08-26 至 2026-08-26 共收录 10 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 5 篇

2608.09696 2026-08-26 cs.AI 版本更新 93%

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

模型发现智能体:用于数据高效发现机制世界模型的大语言模型辅助贝叶斯实验设计

Kevin Murphy

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究提出模型发现智能体(MDA),结合LLM与贝叶斯机制,在少量干预下发现机制世界模型,在三类基准上实现数据高效模型学习与可靠干预预测的SOTA性能。

Comments v4: Major update! Fixed a leak in the prompts for physics and chemistry benchmarks and re-ran experiments (fortunately results did not change much), added Boxing Gym benchmark (requires generating NumPyro code), significantly simplified the figures and evaluation protocol, reframed the narrative around SMC^3, polished the presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21114 2026-08-26 cs.CV cs.AI 版本更新 88%

CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents

CIVA:面向视觉世界模型智能体的评论者诱导价值子空间攻击

Jiancheng Wang, Mingli Zhu, Tong Zhang, Jiaqi Ruan, Wei Wang, Siyuan Liang, Dacheng Tao

专题命中 通用世界模型 :world-model(title,abstract);world-model(title,abstract);分类 cs.AI、cs.CV

AI总结 该研究针对视觉世界模型智能体提出CIVA攻击方法,通过提取价值子空间优化扰动,在多个基准任务上优于现有方法,实现了低时间变化下的显著奖励下降。

Comments Includes supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13554 2026-08-26 cs.ET cs.IT cs.RO 版本更新 81%

Observability Engineering: From Measurement to Information Generation in Active Sensing Systems

从快照感知到持久电磁世界建模:一种面向ISAC的生成空间视角

Pin-Han Ho, Limei Peng

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出一种基于生成空间的毫米波感知框架,通过低维激励空间实现灵活、可扩展的持续电磁世界建模,避免了快照感知的局限性。

Comments 7 pages, 6 figures/tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12144 2026-08-26 cs.CV cs.RO eess.IV 版本更新 71%

O3N: Omnidirectional Open-Vocabulary Occupancy Prediction for Embodied Intelligent Robotics

O3N: 全向开放词汇占用预测

Mengfei Duan, Hao Shi, Fei Teng, Guoqiang Zhao, Yuheng Zhang, Zhiyong Li, Kailun Yang

机构 * Hunan University(湖南大学) Zhejiang University(浙江大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV、cs.RO

AI总结 O3N通过全向感知和开放词汇方法,实现三维空间的连续表示和长距离上下文建模,提升具身智能在开放世界中的感知与建模能力。

Comments The source code will be made publicly available at this https URL (https://github.com/MengfeiD/O3N)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23383 2026-08-26 cs.CV 版本更新 69%

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

面向持续故事与交互世界的长时序视听生成

Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang

机构 * Joy Future Academy, JD(京东探索研究院)

专题命中 通用世界模型 :world-model(abstract);world-model(abstract);分类 cs.CV

AI总结 研究针对长时序视听生成的需求,提出JoyAI-Echo-1.5系统,含长视频与世界模型变体,通过专用技术实现跨镜头一致性等性能,在相关基准上取得领先结果,为生成连贯内容提供基础。

Comments Project page: this https URL (https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身与机器人 2 篇

2608.01381 2026-08-26 cs.RO 版本更新 88%

DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation

DreamTrajectory:面向移动操作的、结合世界模型对齐的轨迹引导动作生成方法

Zheng Yang, Wenjie Zhang, Xiangyu Chen, Wenxuan Song, Xianpeng Wang, Yihang Kang, Jiawen Wen, Wen Chen, Lujia Wang, Renjing Xu, Haoang Li, Xiaowen Chu

专题命中 具身与机器人 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 DreamTrajectory是面向移动操作的轨迹引导框架,通过联合预测末端执行器轨迹与全身动作块、结合轨迹世界模型的测试时细化,在MS-HAB及真实任务中大幅提升操作成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22067 2026-08-26 cs.RO cs.AI cs.CV cs.LG 版本更新 75%

Inferring Action from Future Latent State for Robotic Manipulation

从未来隐状态推断机器人操纵动作

Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Peilin Huang, Han Fu, Zhuo Li, Xiaoxue Ren

机构 * DeepLeap Research(DeepLeap研究院)

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG、cs.CV

AI总结 本文提出无需视频生成的机器人操纵模型DELE-w0.5,通过从捕获动作相关物理结果的未来隐状态推断动作,在4项长程操纵任务的480次试验中,其性能优于最强基线47.5和30.7个百分点,实现最优表现。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 自动驾驶 2 篇

2605.10426 2026-08-26 cs.CV cs.AI 版本更新 85%

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving

CoWorld-VLA:面向自动驾驶的多专家世界模型中的思考

Minqing Huang, Yujiao Xiang, Zihan Liang, Jiajie Huang, Jingqi Wang, Yuheng Zhou, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang, Gong Che

机构 * Afari Intelligent Drive(Afari智能驾驶公司) University of Electronic Science and Technology of China(电子科技大学) Shanghai Jiao Tong University(上海交通大学) Beijing University Of Posts and Telecommunications(北京邮电大学) Tianjin University(天津大学)

专题命中 自动驾驶 :world model(title);world model(title);分类 cs.AI、cs.CV

AI总结 本文提出CoWorld-VLA,通过多专家世界推理框架,利用显式条件指导动作规划,提升自动驾驶的场景生成与路径规划性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23486 2026-08-26 cs.CV cs.RO 版本更新 71%

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

GeoWAM:面向自动驾驶的视觉几何世界动作模型

Yiren Lu, Xin Ye, Jiaming Liu, Philip Jacobson, Jin Yao, Yi-chung Chen, Liam Merino, Dhruva Dixith Kurra, Min Cai, Tom Lampo, Yu Yin, Danhua Guo, Burhan Yaman

机构 * Uber AV Labs(优步自动驾驶实验室) Case Western Reserve University(凯斯西储大学)

专题命中 自动驾驶 :world model(abstract);world model(abstract);分类 cs.CV、cs.RO

AI总结 该研究提出 GeoWAM,一种用于自动驾驶的视觉几何世界动作模型,通过预训练预测未来场景几何来学习动力学,经评估其生成的驾驶策略比图像基方案更强,确立未来几何预测为自动驾驶有效预训练目标。

Comments Project page: this https URL (https://yiren-lu.com/project_pages/geowam/)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 仿真与规划 1 篇

2603.14603 2026-08-26 cs.RO 版本更新 68%

Latent Dynamics-Aware OOD Monitoring for Trajectory Prediction with Provable Guarantees

隐式动态感知的领域外监测用于轨迹预测的可证明保障

Tongfei Guo, Lili Su

专题命中 仿真与规划 :latent dynamics(title);分类 cs.RO

AI总结 本文提出基于快速突变点检测的轨迹预测领域外监测方法,通过隐马尔可夫模型建模预测误差演化,实现无需显式知识的领域外检测并保证延迟和误报率的可证明保障。

Comments Accepted by 2026 IEEE International Conference on Automation Science and Engineering (CASE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏