arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-01-28 至 2026-01-28 共收录 12 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 9 篇

2601.19834 2026-01-28 cs.AI 94%

Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models

视觉生成通过多模态世界模型解锁类人推理

Jialong Wu, Xiaoying Zhang, Hongyi Yuan, Xiangcheng Zhang, Tianhao Huang, Changjing He, Chaoyi Deng, Renrui Zhang, Youbin Wu, Mingsheng Long

机构 * Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出视觉生成在特定任务中优于纯语言推理,通过构建VisWorld-Eval评估套件验证了多模态世界模型提升类人推理的能力。

Comments Project page: https://thuml.github.io/Reasoning-Visual-World

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19336 2026-01-28 cs.LG cs.AI 91%

From Observations to Events: Event-Aware World Model for Reinforcement Learning

从观测到事件:用于强化学习的事件感知世界模型

Zhao-Han Peng, Shaohui Li, Zhi Li, Shulan Ruan, Yu Liu, You He

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出事件感知世界模型(EAWM),通过学习事件感知的表示来提升强化学习的样本效率,实验表明其在多个基准上均取得显著性能提升。

Comments 43 pages, accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08931 2026-01-28 cs.CV cs.AI cs.LG 91%

Astra: General Interactive World Model with Autoregressive Denoising

Astra:通用交互世界模型与自回归去噪

Yixuan Zhu, Jiaqi Feng, Wenzhao Zheng, Yuan Gao, Xin Tao, Pengfei Wan, Jie Zhou, Jiwen Lu

机构 * Tsinghua University(清华大学) Kuaishou Technology(快手科技)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 Astra提出了一种通用交互世界模型,通过自回归去噪和动作感知适配器实现长周期视频预测与多样化交互。

Comments Accepted in ICLR 2026. Code is available at: https://github.com/EternalEvan/Astra

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19484 2026-01-28 cs.CV 81%

Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes

动态世界,动态人类:生成动态场景中的虚拟人类-场景交互动作

Yin Wang, Zhiying Leng, Haitian Liu, Frederick W. B. Li, Mu Li, Xiaohui Liang

机构 * Beihang University(北航大学) University of Durham(达勒姆大学) State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室) Zhongguancun Laboratory(中关村实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出Dyn-HSI架构,通过动态视觉导航、分层经验记忆和人-场景交互扩散模型,生成高质量的动态场景中人-场景交互动作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13832 2026-01-28 astro-ph.EP astro-ph.SR 80%

TOI-333b: A Neptune Desert planet around a F7V star

TOI-333b:围绕F7V恒星的海王星荒漠行星

Douglas R. Alves, James S. Jenkins, José I. Vinés, Maximilano Moyano, David R. Anderson, Christian Magliano, Giovanni Covone, Keivan G. Stassun, Abderahmane Soubkiou, Edward Gillen, Matthew P. Battley, Alexander Hughes, David J. Armstrong, Suman Saha, Faith Hawthorn, Peter J. Wheatley, Karen A. Collins, Richard P. Schwarz, Gregor Srdoc, Ioannis Apergis, Tafadzwa Zivave, Monika Lendl, Benjamin M. Tofflemire, John P. Doty, Christina Hedges, Ismael Mireles, Matthew R. Burleigh, Alicia Kendall, George T. Harvey, Michael R. Goad, Sarah L. Casewell, Troy Edkins

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 TOI-333b是围绕F7V恒星的海王星荒漠行星,其质量、半径和密度等特征揭示了其可能的内部组成和演化过程,为研究此类行星在热恒星周围的演化提供了独特实验室。

Comments 21 pages, 19 figures, 7 tables, accepted for publication in A&A

Journal ref A&A 705, A210 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19819 2026-01-28 gr-qc 80%

Comment on "Multidimensional arrow of time" (arXiv:2601.14134)

对“多维时间之箭”的评论(arXiv:2601.14134)

Andrei Galiautdinov

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文针对Rubin关于时间之箭源于额外维度体积增长的观点,提出"形状动态时间之箭"概念,利用Perelman熵的性质解决体积增长与引力常数观测稳定性之间的矛盾。

Comments Comment on arXiv:2601.14134; 6 pages, no figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19752 2026-01-28 cs.AI 69%

Agentic Design Patterns: A System-Theoretic Framework

代理设计模式:一种系统理论框架

Minh-Dung Dao, Quy Minh Le, Hoang Thanh Lam, Duc-Trong Le, Quoc-Viet Pham, Barry O'Sullivan, Hoang D. Nguyen

机构 * University College Cork Vietnam National University(越南国家大学) IBM Research Ireland(IBM爱尔兰研究) Trinity College Dublin(都柏林三一学院)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 本文提出了一种系统理论框架,通过分解代理AI系统为五个核心功能子系统,提出12种代理设计模式,以解决代理设计中的重复问题,提升自主系统的模块化和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19712 2026-01-28 cs.SD cs.MM 65%

Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling

具有视觉-语言先验和3D声学环境建模的物理感知新视角声音合成

Congyi Fan, Jian Guan, Youtian Lin, Dongli Xu, Tong Ye, Qiaoxi Zhu, Pengming Feng, Wenwu Wang

机构 * Group of Intelligent Signal Processing(智能信号处理组) Harbin Engineering University(哈尔滨工程大学) School of Intelligence Science and Technology(智能科学与技术学院) Nanjing University(南京大学) Processing Speech and Images(语音与图像处理) KU Leuven(库尔勒文大学) Acoustics Lab(声学实验室) University of Technology Sydney(悉尼技术大学) State Key Laboratory of Space Information System and Integrated Application(空间信息系统与集成应用国家重点实验室) Centre for Vision Speech and Signal Processing(视觉语音与信号处理中心) University of Surrey(萨里大学)

专题命中 通用世界模型 :environment model(title)

AI总结 Phys-NVAS通过整合视觉-语言语义先验与3D声学环境建模,实现了具有物理感知的新视角声音合成,提升了声音的真实感和物理一致性。

Comments ICASSP 2026 Accept, Project page: https://physnvas.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06652 2026-01-28 cs.LG cs.AI 50%

Adaptive Test-Time Training for Predicting Need for Invasive Mechanical Ventilation in Multi-Center Cohorts

自适应测试时训练用于多中心队列中预测有创机械通气需求

Xiaolei Lu, Shamim Nemati

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 本文提出自适应测试时训练框架,通过信息论界限和自监督学习提升多中心ICU中IMV预测的泛化性能和鲁棒性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身与机器人 1 篇

2601.19839 2026-01-28 cs.RO cs.AI cs.HC 71%

HARMONI: Multimodal Personalization of Multi-User Human-Robot Interactions with LLMs

HARMONI: 基于大语言模型的多用户人机交互多模态个性化

Jeanne Malécot, Hamed Rahimi, Jeanne Cattoni, Marie Samson, Mouad Abrini, Mahdi Khoramshahi, Maribel Pino, Mohamed Chetouani

机构 * Institut Curie, Université Paris-Saclay(巴黎-萨克勒大学Curie研究所) Institute of Intelligent Systems and Robotics (ISIR), Sorbonne University(索邦大学智能系统与机器人研究所) Assistance Publique – Hôpitaux de Paris (AP-HP), Université Paris Cité(巴黎公共医院(AP-HP)与巴黎城市大学)

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.AI、cs.RO

AI总结 HARMONI通过多模态个性化框架,利用大语言模型提升多用户人机交互的持续个性化和动态适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 模型式强化学习 2 篇

2410.11234 2026-01-28 cs.LG cs.AI 88%

Bayes Adaptive Monte Carlo Tree Search for Offline Model-based Reinforcement Learning

贝叶斯自适应蒙特卡洛树搜索用于离线模型驱动强化学习

Jiayu Chen, Le Xu, Wentse Chen, Jeff Schneider

机构 * The University of Hong Kong(香港大学) Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 模型式强化学习 :model-based reinforcement learning(title,abstract);world model(abstract);world models(abstract);world model(abstract)

AI总结 本文提出贝叶斯自适应蒙特卡洛树搜索算法,用于提升离线模型驱动强化学习的性能,显著优于现有方法。

Comments This paper is accepted in ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19551 2026-01-28 cs.LG cs.AI 56%

Scale-Consistent State-Space Dynamics via Fractal of Stationary Transformations

通过稳定变换的分形实现尺度一致的状态空间动力学

Geunhyeok Yu, Hyoseok Hwang

机构 * Department of Software Convergence, Kyung Hee University(软件融合系,庆熙大学)

专题命中 模型式强化学习 :latent dynamics(abstract);分类 cs.AI、cs.LG

AI总结 本文提出FROST方法,通过分形归纳偏置实现状态空间模型的尺度一致潜在动力学,提升模型的自适应效率和稳定性。

Comments 8 pages (excluding 2 pages of references), 3 tables, 2 figures. Appendix: 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏