arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-03-30 至 2026-03-30 共收录 10 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 6 篇

2511.18746 2026-03-30 cs.CV cs.AI 96%

Any4D: Open-Prompt 4D Generation from Natural Language and Images

Any4D: 从自然语言和图像生成开放提示的4D生成

Hao Li, Qiao Sun

专题命中 通用世界模型 :world model(summary_cn,abstract);world models(summary_cn,abstract);embodied world model(summary_cn,abstract);world model(summary_cn,abstract)

AI总结 本文提出Primitive Embodied World Models,通过限制视频生成时间范围,实现语言与视觉表示的细粒度对齐,降低学习复杂度,提升数据效率,并减少推理延迟,支持复杂任务的组合泛化。

Comments The authors identified issues in the 4D generation pipeline and evaluation that affect result validity. To ensure scientific accuracy, we will revise the methodology and experiments thoroughly before resubmitting. This version should not be cited or relied upon

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25981 2026-03-30 cs.RO cs.AI cs.CL 90%

Policy-Guided World Model Planning for Language-Conditioned Visual Navigation

基于策略引导的世界模型规划用于语言条件的视觉导航

Amirhosein Chahe, Lifeng Zhou

机构 * Drexel University(德雷塞尔大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出PiJEPA框架,结合学习导航策略与潜在世界模型规划,通过微调Octo策略和冻结预训练视觉编码器,提升语言指导下的视觉导航准确性与指令遵循性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14375 2026-03-30 cs.CV cs.AI 82%

The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics

运动的脉搏:从视觉动态测量物理帧率

Xiangbo Gao, Mingyang Wu, Siyuan Yang, Jiongze Yu, Pardis Taghavi, Fangzhou Lin, Zhengzhong Tu

机构 * Texas A\&M University

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出Visual Chronometer,通过视觉动态直接恢复物理帧率,解决视频生成中因训练数据帧率不一致导致的物理运动速度模糊问题,提升生成视频的自然度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25887 2026-03-30 cs.CV 81%

World Reasoning Arena

世界推理竞技场

PAN Team, Qiyue Gao, Kun Zhou, Jiannan Xiang, Zihan Liu, Dequan Yang, Junrong Chen, Arif Ahmad, Cong Zeng, Ganesh Bannur, Xinqi Huang, Zheqi Liu, Yi Gu, Yichi Yang, Guangyi Liu, Zhiting Hu, Zhengzhong Liu, Eric Xing

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出WR-Arena基准,评估世界模型在动作模拟、长期预测和推理规划三方面的能力,揭示当前模型与人类水平推理的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03399 2026-03-30 cs.LG 58%

Robust Predictive Modeling Under Unseen Data Distribution Shifts: A Methodological Commentary

在未见数据分布偏移下实现稳健的预测建模:一种方法论评论

Hanyu Duan, Yi Yang, Ahmed Abbasi, Kar Yan Tam

机构 * Hong Kong University of Science and Technology(香港科技大学) University of Notre Dame(圣母大学)

专题命中 通用世界模型 :predictive model(title,abstract);分类 cs.LG;predictive models(abstract)

AI总结 本文通过实际客户流失案例,指出训练与测试数据分布不一致的问题,并探讨领域泛化方法,提出采用不确定性意识的预测建模方法以提升模型鲁棒性。

Comments Forthcoming in Information Systems Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25901 2026-03-30 cs.LG cs.AI cs.CV 50%

Decoding Defensive Coverage Responsibilities in American Football Using Factorized Attention Based Transformer Models

利用基于因子化注意力的变换器模型解码美式足球防守覆盖责任

Kevin Song, Evan Diewald, Ornob Siddiquee, Chris Boomhower, Keegan Abdoo, Mike Band, Amy Lee

机构 * Amazon Web Services, Seattle, WA, USA(亚马逊云服务(AWS),美国华盛顿州西雅图) National Football League, New York, NY, USA(美国国家橄榄球联盟(NFL),美国纽约州纽约市)

专题命中 通用世界模型 :分类 cs.AI、cs.LG、cs.CV;predictive model(abstract)

AI总结 本文提出基于因子化注意力的变换器模型,用于预测美式足球比赛中防守球员的覆盖分配、接球手与防守球员的配对及每次传球的受攻防守球员,实现了对球员责任的动态预测。

Comments 19 pages, 8 figures, ISACE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身与机器人 1 篇

2603.23376 2026-03-30 cs.CV cs.RO 82%

ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment

ABot-PhysWorld:用于机器人操作的交互世界基础模型,具有物理对齐

Yuzhi Chen, Ronghan Chen, Dongjie Huo, Yandan Yang, Dekang Qi, Haoyun Liu, Tong Lin, Shuang Zeng, Junjin Xiao, Xinyuan Chang, Feng Xiong, Xing Wei, Zhiheng Ma, Mu Xu

专题命中 具身与机器人 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 ABot-PhysWorld通过物理对齐的交互世界模型,生成物理合理且可控的视频,解决了传统方法在物理仿真中的不足,通过新框架和基准测试提升了性能。

Comments Code: https://github.com/amap-cvlab/ABot-PhysWorld.git

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 自动驾驶 2 篇

2509.15219 2026-03-30 cs.CV cs.LG cs.MA cs.MM cs.RO 83%

Out-of-Sight Embodied Agents: Multimodal Tracking, Sensor Fusion, and Trajectory Forecasting

视线外的具身智能体:多模态跟踪、传感器融合与轨迹预测

Haichao Zhang, Yi Xu, Yun Fu

机构 * Department of Electrical and Computer Engineering, Northeastern University(东北大学电气与计算机工程系) Khoury College of Computer Sciences, Northeastern University(东北大学库里计算机科学学院)

专题命中 自动驾驶 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出改进的视线外轨迹预测方法,通过视觉定位去噪模块提升轨迹预测性能,实验表明在Vi-Fi和JRDB数据集上达到最优效果,为自动驾驶、机器人和监控提供新方向。

Comments Published in IEEE Transactions on Pattern Analysis and Machine Intelligence (Early Access), pp. 1-14, March 23, 2026

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24989 2026-03-30 cs.RO cs.AI 53%

Learning Rollout from Sampling:An R1-Style Tokenized Traffic Simulation Model

从采样中学习:一种R1风格的分词交通仿真模型

Ziyan Wang, Peng Chen, Ding Li, Chiwei Li, Qichao Zhang, Zhongpu Xia, Guizhen Yu

机构 * State Key Laboratory of Intelligent Transportation System, Key Laboratory of Autonomous Transportation Technology for Special Vehicles, Ministry of Industry and Information Technology, School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院,智能交通系统国家重点实验室,特种车辆自主运输技术重点实验室(工业和信息化部)) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所,多模态人工智能系统国家重点实验室)

专题命中 自动驾驶 :simulation model(title);分类 cs.AI、cs.RO

AI总结 本文提出R1Sim模型,通过分词交通仿真结合熵引导采样和GRPO优化,实现探索与利用的平衡,生成真实安全的多智能体行为。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 仿真与规划 1 篇

2603.23610 2026-03-30 cs.AI 56%

Environment Maps: Structured Environmental Representations for Long-Horizon Agents

环境地图:面向长周期智能体的结构化环境表示

Yenchia Feng, Chirag Sharma, Karime Maamari

机构 * Distyl AI

专题命中 仿真与规划 :world model(comments);world models(comments);分类 cs.AI;world model(comments)

AI总结 本文提出环境地图,通过整合异构证据构建结构化图表示,提升长周期任务的鲁棒性。在WebArena基准上,环境地图使智能体成功率提升至28.2%,优于基线方法。

Comments 9 pages, 5 figures, accepted to ICLR 2026 the 2nd Workshop on World Models; updated formatting issue

详情

展开后加载摘要…

URL PDF HTML 收藏