arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 3110 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 3110 篇

2606.25473 2026-06-25 cs.CV cs.LG 新提交 81%

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models

Causal-rCM:一种用于流式视频生成和交互式世界模型中自回归扩散蒸馏的统一教师强制与自我强制开放配方

Kaiwen Zheng, Guande He, Min Zhao, Jintao Zhang, Huayu Chen, Jianfei Chen, Chen-Hsuan Lin, Ming-Yu Liu, Jun Zhu, Qianli Ma

机构 * Tsinghua University(清华大学) UT Austin(德克萨斯大学奥斯汀分校) NVIDIA(英伟达)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV、cs.LG

AI总结 提出Causal-rCM框架,将扩散蒸馏扩展到自回归视频扩散,通过教师强制(前向散度)与自我强制(反向散度)互补,实现流式视频生成和交互式世界模型,在帧级和块级设置中达到最先进性能。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24946 2026-06-25 cs.LG cs.RO 新提交 81%

Conformal Orbit-Valid Trust Horizons for Equivariant World Models

等变世界模型的共形轨道有效信任视界

Hongbo Wang

机构 * Department of Mathematics, Stony Brook University(石溪大学数学系)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.LG

AI总结 针对具有已知群对称性的潜在世界模型,提出一种共形校准的信任视界认证方法,利用等变性实现轨道常数认证,实验验证了其保守性和非空性。

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17796 2026-06-25 cs.CV cs.AI 版本更新 81%

CustomX: Unified Character, Action, and Scene Customization in Video World Models

CustomX: 视频世界模型中的统一角色、动作与场景定制

Yitong Wang, Fangyun Wei, Hongyang Zhang, Bo Dai, Yan Lu

机构 * Fudan University(复旦大学) Microsoft Research(微软研究院) University of Waterloo(滑铁卢大学) The University of Hong Kong(香港大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 提出CustomX,结合静态世界生成与可控实体模型,支持用户指定角色在3D场景中执行开放动作,通过条件自回归视频生成保持视觉保真度。

Comments Accepted to ECCV 2026. Project page: https://snowflakewang.github.io/CustomX_Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24089 2026-06-24 cs.RO cs.AI 新提交 81%

DynaWM: Dynamics-Aware Distillation with World Model and Momentum Targets for Smooth Locomotion over Continuous Stairs

DynaWM: 基于世界模型和动量目标的动力学感知蒸馏实现连续楼梯上的平滑运动

Haidong Hou, Zhangguo Yu, Hengbo Qi, Jianlin Zhang

机构 * School of Mechatronical Engineering, Beijing Institute of Technology(北京理工大学机电学院)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.AI

AI总结 提出DynaWM框架,通过世界模型正则化增强地形编码,并利用动量目标编码器稳定知识蒸馏,使双足轮式机器人在连续楼梯上实现高适应性和平滑运动。

Comments Comments: 8 pages, 7 figures, accepted by IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

Journal ref IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS),2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22449 2026-06-23 cs.AI cs.RO 新提交 81%

Self-Evolving Cognitive Framework via Causal World Modeling for Embodied Scientific Intelligence

基于因果世界建模的自进化认知框架用于具身科学智能

Yi Yu, Tetsunari Inamura

机构 * Graduate School of Advanced Science and Engineering, Hiroshima University(广岛大学先进科学与工程研究生院) Advanced Intelligence and Robotics Research Center, Brain Science Institute, Tamagawa University(玉川大学脑科学研究所先进智能与机器人研究中心)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.AI

AI总结 提出一种自进化认知框架,通过因果世界建模、干预驱动推理和持续认知精炼,使具身智能体在交互中不断构建和修正内部因果模型,实现从预测智能到认知智能的转变。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21775 2026-06-23 cs.LG cs.AI 新提交 81%

Beyond the Next Step: Variable-Length Latent World Models for Long-Horizon Planning

超越下一步:用于长时程规划的变长潜世界模型

Tianqi Du, Qi Zhang, Yifei Wang, Yisen Wang

机构 * State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院通用人工智能国家重点实验室) Amazon AGI SF Lab(亚马逊AGI旧金山实验室) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出变长潜世界模型(VLWM),通过学习变长动作序列的条件潜状态预测,解决递归一步预测在长时程规划中的累积误差问题,结合课程训练策略,在长时程控制任务上平均提升13%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21173 2026-06-23 cs.LG cs.AI 新提交 81%

Inverting the Bellman Equation: From $Q$-Values to World Models

逆推贝尔曼方程:从 $Q$ 值到世界模型

Alistair Letcher, Mattie Fellows, Alexander D. Goldie, Jonathan Richens, Jakob N. Foerster, Oliver Richardson

机构 * FLAIR, University of Oxford(FLAIR,牛津大学) Google DeepMind(谷歌DeepMind) Mila, University of Montreal(Mila,蒙特利尔大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文证明基于值的智能体在丰富奖励函数上训练时隐式编码世界模型,提出 $P$-learning 从 $Q$ 值提取模型,并给出编码真实转移核的充分条件,实验验证了隐式模型的准确性和泛化能力。

Comments 48 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20104 2026-06-19 cs.LG cs.AI 新提交 81%

Sensorimotor World Models: Perception for Action via Inverse Dynamics

传感器运动世界模型:通过逆动力学实现面向行动感知

Petr Ivashkov, Randall Balestriero, Bernhard Schölkopf

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) Department of Computer Science, Brown University(布朗大学计算机科学系) ELLIS Institute(ELLIS研究所) ETH Zürich(苏黎世联邦理工学院)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出传感器运动世界模型(SMWM),通过逆动力学正则化端到端训练潜空间世界模型,防止表示崩溃并学习与行动对齐的紧凑表示,在2D和3D控制任务中实现竞争性规划性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18688 2026-06-18 cs.LG cs.AI 新提交 81%

Dual-Channel Grounded World Modeling (DCGWM): Structural Prevention of Objective Interference Collapse via Heterogeneous External Grounding with Inward-Only Gradient Flow

双通道接地世界建模 (DCGWM):通过异构外部接地与内向梯度流结构性防止目标干扰崩溃

Akshay Hazare

机构 * Independent Researcher(独立研究者)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出双通道接地世界建模(DCGWM),通过分区潜空间和内向梯度流,结构性防止联合嵌入预测架构中多目标接地导致的目标干扰崩溃。

Comments Position paper. Experimental validation in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18582 2026-06-18 cs.CV cs.RO eess.IV 新提交 81%

Technical Report for ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Leveraging DINOv3 for Robust Outdoor Scene Understanding in Field Robotics

ICRA 2026 GOOSE 2D细粒度语义分割挑战赛技术报告:利用DINOv3实现野外机器人中的鲁棒户外场景理解

Jaeil Park, Hyobin Choi, Sangjin Lee, Hyungtae Lim, Sung-Hoon Yoon

机构 * Daegu Gyeongbuk Institute of Science and Technology (DGIST)(大邱庆北科学技术院) Massachusetts Institute of Technology (MIT)(麻省理工学院)

专题命中 具身推理 :robotics(title,abstract);分类 cs.RO、cs.CV

AI总结 提出一种结合DINOv3自监督骨干、ViT-Adapter和Mask2Former解码器的网络设计,以及多尺度测试增强和模型集成的推理策略,在64类细粒度越野语义分割挑战中取得第一名,复合得分76.57%。

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17536 2026-06-17 cs.CV cs.AI 新提交 81%

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation

OmniDrive: 一种由LLM编排的多智能体世界模型,用于多视角驾驶视频生成的统一潜在协同压缩

Zijie Meng, Yufei Liu, Chengqian Ma, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Shuqin Chen, Weichen Xu, Jiquan Yuan, Miao Zhang

机构 * Peking University(北京大学) Xiamen University(厦门大学) Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) National Taiwan University(国立台湾大学) Wuhan University(武汉大学) Wuhan University of Technology(武汉理工大学) Tsinghua University(清华大学) Jimei University(集美大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 提出DRIVE-CHOREO,一种由LLM编排的多智能体世界模型,通过三个Qwen2.5-VL智能体协同生成位置感知的潜在序列,并利用视图-时间置换与3D VAE协同压缩,实现可控多视角视频生成,在nuScenes上达到SOTA多视角一致性和BEV mAP 21.6。

Comments 24 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18701 2026-06-17 cs.LG cs.AI stat.ML 版本更新 81%

Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training

Curiosity-Critic:累积预测误差改进作为世界模型训练的可处理内在奖励

Vin Bhaskara, Haicheng Wang

机构 * Department of Computer Science, University of Toronto, Toronto, Canada(多伦多大学计算机科学系)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出Curiosity-Critic方法,通过可处理的每步替代项(当前预测误差与渐近误差基线的差值)作为内在奖励,利用共训练的评论家在线估计误差基线,有效分离可约与不可约预测误差,在随机网格世界实验中优于现有方法。

Comments Accepted to ICML 2026 Workshop on Epistemic Intelligence in Machine Learning (EIML@ICML 2026). Code: https://github.com/vinbhaskara/Curiosity-Critic

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16076 2026-06-16 cs.LG cs.AI cs.GT 新提交 81%

Phys-JEPA: Physics-Informed Latent World Models for Multivariate Time-Series Forecasting

Phys-JEPA:面向多变量时间序列预测的物理信息潜在世界模型

Weizhi Nie, Weichao Liu, Honglin Guo, Yuting Su

机构 * Tianjin University(天津大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出Phys-JEPA架构,将物理一致性约束引入潜在状态和状态转移,分解预测状态为物理和残差分量,在气候、交通、电力数据集上提升预测精度。

Comments Submitted to arXiv as a preliminary manuscript. 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15160 2026-06-16 cs.CV cs.LG 新提交 81%

DLWM: Diverse Latent World Models for Efficient Multimodal Reasoning

DLWM: 多样化潜在世界模型用于高效多模态推理

David Huang, Lianlei Shan

机构 * University of Toronto(多伦多大学) Tsinghua University(清华大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV、cs.LG

AI总结 提出DLWM框架,结合潜在空间推理与强化学习,通过多样化潜在假设和资源感知策略提升多模态推理效率,准确率提升2-5%,内存减少24%。

Comments Preprint. 9 pages main text, 15 pages total including appendix, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14934 2026-06-16 cs.LG cs.AI 新提交 81%

Separable Neural Architectures as Physical World Models: from Mathematical Theory to Applications

可分离神经架构作为物理世界模型:从数学理论到应用

Reza T Batley, Andrew Kichline, Sourav Saha

机构 * Kevin T. Crofton Department of Aerospace and Ocean Engineering, Virginia Polytechnic Institute and State University(弗吉尼亚理工大学凯文·T·克罗夫顿航空航天与海洋工程系)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出可分离神经架构(SNA),结合神经逼近与张量分解,通过变分框架求解偏微分方程,实现高维问题代数级缩放,并在工程案例中取得显著加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21577 2026-06-16 cs.CL cs.AI cs.LG stat.ML 版本更新 81%

A Unified Definition of Hallucination: It's The World Model, Stupid!

幻觉的统一定义:是世界模型的问题,笨蛋!

Emmy Liu, Varun Gangal, Chelsea Zou, Michael Yu, Xiaoqi Huang, Alex Chang, Zhuofu Tao, Karan Singh, Sachin Kumar, Steven Y. Feng

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出幻觉的统一定义,即用户可观察到的错误内部世界建模,并连接至HalluWorld基准测试,以区分真实幻觉与规划或奖励错误。

Comments ICML 2026. HalluWorld benchmark at https://github.com/DegenAI-Labs/HalluWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01801 2026-06-15 cs.CV cs.AI 版本更新 81%

Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention

快速自回归视频扩散与世界模型:基于时间缓存压缩与稀疏注意力

Dvir Samuel, Issar Tzachor, Matan Levy, Michael Green, Gal Chechik, Rami Ben-Ari

机构 * Hebrew University of Jerusalem(特拉维夫大学) Google Research(谷歌研究)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 提出FAST-AR框架,通过TempCache压缩KV缓存、AnnCA加速交叉注意力、AnnSA稀疏化自注意力,实现自回归视频扩散模型5-10倍加速,同时保持视觉质量并稳定GPU内存使用。

Comments Accepted to ICML 2026. Project Page: https://dvirsamuel.github.io/fast-auto-regressive-video/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11395 2026-06-12 cs.LG cs.AI 版本更新 81%

ARROW: Augmented Replay for RObust World models

ARROW:增强重放用于鲁棒世界模型

Abdulaziz Alyahya, Abdallah Al Siyabi, Markus R. Ernst, Luke Yang, Levin Kuhlmann, Gideon Kowadlo

机构 * Imam Mohammad Ibn Saud Islamic University (IMSIU)(伊玛姆·穆罕默德·本·沙特伊斯兰大学) Monash University(莫纳什大学) University of New South Wales, Sydney(新南威尔士大学,悉尼) Cerenaut

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出ARROW算法,一种基于模型的持续强化学习方法,通过高效的重放缓冲区减少灾难性遗忘,提升在无共享结构任务和有共享结构任务中的表现。

Comments 36 pages and 11 figures (includes Appendix)

Journal ref Transactions on Machine Learning Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16713 2026-06-12 cs.CV cs.AI 版本更新 81%

GeoWorld-VLM: Geometry from World Models for Vision-Language Models

GeoWorld-VLM:从世界模型中获取几何结构用于视觉-语言模型

Renjie Gu, Kaichen Zhou, Yan Luo, Mengyu Wang

机构 * Harvard AI and Robotics Lab(哈佛人工智能与机器人实验室) Kempner Institute for the Study of Natural and Artificial Intelligence(凯普纳自然与人工智能研究 institute) Harvard University(哈佛大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 GeoWorld-VLM通过将冻结的摄像机条件视频世界模型的几何结构转移到视觉-语言模型中,提升空间关系推理能力,实验显示在两个不同架构上均提升了约4%的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10620 2026-06-10 cs.CV cs.AI 新提交 81%

Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency

图像模型能想象时间吗?ImageTime:通过时空一致性探究视觉世界建模的新基准

Xinrui Wu, Lichen Huang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 提出ImageTime基准,通过四关键帧协议(初始状态、动作开始、过渡状态、最终状态)评估图像生成模型在时空一致性上的表现,揭示模型在维持连贯视觉世界状态方面的能力与不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09803 2026-06-09 cs.CV cs.GR cs.LG 新提交 81%

Echo-Memory: A Controlled Study of Memory in Action World Models

Echo-Memory:动作世界模型中记忆的受控研究

Wayne King, Zeyue Xue, Yuxuan Bian, Jie Huang, Haoran Li, Yaowei Li, Yaofeng Su, Yuming Li, Haoyu Wang, Shiyi Zhang, Songchun Zhang, Yuwei Niu, Sihan Xu, Junhao Zhuang, Haoyang Huang, Nan Duan

机构 * Joy Future Academy(京东探索研究院)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV、cs.LG

AI总结 提出Echo-Memory框架,通过控制变量法研究动作条件世界模型中的记忆机制,发现原始上下文容量和块状状态空间递归对开放域返回任务至关重要。

Comments 9 figures and 28 pages, Code at \href{https://github.com/Echo-Team-Joy-Future-Academy-JD/Echo-Memory}{this URL}

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07974 2026-06-09 cs.RO cs.AI 新提交 81%

PRISM: PRior-guided Imagination Sampling in world Models

PRISM:世界模型中基于先验引导的想象采样

Yuhai Wang, Jiawei Xia, Rongxuan Zhou, Xiao Hu, Yongliang Shi, Jing Du, Yang Ye

机构 * Northeastern University(东北大学) University of California, Berkeley(加州大学伯克利分校) Qiyuan Lab(启元实验室) University of Florida(佛罗里达大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.AI

AI总结 提出PRISM框架,通过从世界模型编码器提取状态条件高斯先验,并利用精度加权高斯乘积更新规划器的采样分布,在不增加架构复杂度的情况下显著提升基于模型的连续控制性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06014 2026-06-05 cs.AI cs.RO 81%

PLAN-S: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models

PLAN-S:通过潜在风格动态桥接规划以实现自动驾驶世界模型

Xiaoyun Qiu, Jingtao He, Yijie Chen, Yusong Huang, Haotian Wang, Yixuan Wang, Xinhu Zheng

机构 * Intelligent Transportation Thrust, Systems Hub, and Center of Seamless Connectivity & Connected Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(智能交通 thrust、系统中心及无缝连接与智能连接研究院,香港科学与技术大学(广州))

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.AI

AI总结 提出PLAN-S框架,通过从潜在表示解码风格条件语义成本图,解决自动驾驶中潜在世界模型规划的可控性问题,在nuScenes和NAVSIM上降低了碰撞率并提升了驾驶性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05979 2026-06-05 cs.RO cs.AI 81%

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

世界-语言-动作模型:统一世界建模、语言推理与动作合成

Yi Yang, Zhihong Liu, Siqi Kou, Yiyang Chen, Yanzhe Hu, Jianbo Zhou, Boyuan Zhao, Zhijie Wei, Xiao Xia, Xueqi Li, Pengfei Liu, Zhijie Deng

机构 * SJTU(上海交通大学) SII(上海研究院) HUST(华中科技大学) SCUT(华南理工大学) ECUST(东华大学) SHU(上海大学) NJUPT(南京工业大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.AI

AI总结 提出世界-语言-动作(WLA)模型,通过自回归Transformer联合预测文本子任务、子目标图像和机器人动作,融合世界建模与语言推理能力,实现多任务和长时域任务的最优性能。

Comments 19 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02697 2026-06-04 cs.CV cs.AI 81%

ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling

ShareVerse:面向共享世界建模的多智能体一致视频生成

Jiayi Zhu, Jianing Zhang, Yiying Yang, Wei Cheng, Xiaoyun Yuan

机构 * Shanghai Jiao Tong University China(上海交通大学中国) Fudan University China(复旦大学中国) StepFun China(StepFun中国)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 提出ShareVerse框架,通过构建多智能体交互数据集、空间拼接策略和跨智能体注意力机制,实现多智能体共享世界的一致视频生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03685 2026-06-03 cs.LG cs.AI 81%

A Close Look At World Model Recovery In Supervised Fine-Tuned LLM Planners

监督微调的大语言模型规划器中世界模型恢复的深入探究

Patrick Emami, Nan Qiang, Peter Graf

机构 * National Laboratory of the Rockies(落基山国家实验室)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 通过可解释性实验,研究监督微调如何影响大语言模型在经典规划任务中恢复世界模型的能力,发现微调使模型线性编码动作有效性和状态谓词,且更广泛的状态空间覆盖有助于更准确的世界模型恢复。

Comments 17 pages. Under review at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18690 2026-06-03 q-bio.NC cs.CV cs.LG 81%

Neural Fields as World Models

神经场作为世界模型

Joshua Nunley

机构 * Luddy School of Informatics, Computing, and Engineering, Indiana University, Bloomington(信息学、计算与工程学院,印第安纳大学,布卢明顿) Cognitive Science Program, Indiana University, Bloomington(认知科学项目,印第安纳大学,布卢明顿)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV、cs.LG

AI总结 提出同构世界模型,利用运动门控神经场在空间图中进行物理预测,实现离线任务学习和身体相关表征。

Comments 6 pages, 6 figures. Annual Meeting of the Cognitive Science Society (CogSci 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02388 2026-06-02 cs.LG cs.AI 81%

Policy and World Modeling Co-Training for Language Agents

语言智能体的策略与世界模型协同训练

Ning Lu, Baijiong Lin, Shengcai Liu, Jiahao Wu, Haoze Lv, Yanbin Wei, Lingting Zhu, Shengju Qian, Xin Wang, Ying-Cong Chen, Qi Wang, Ke Tang

机构 * Southern University of Science and Technology(南方科技大学) Hong Kong University of Science and Technology(香港科学大学) Hong Kong University of Science and Technology (Guangzhou)(香港科学大学(广州)) Hong Kong Polytechnic University(香港理工大学) LIGHTSPEED

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出PaW框架,通过在强化学习过程中添加辅助世界模型监督,无需改变推理范式,提升语言智能体在多个任务上的性能。

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18803 2026-06-01 cs.LG cs.AI 81%

PROWL: Prioritized Regret-Driven Optimization for World Model Learning

PROWL: 基于优先遗憾驱动的世界模型学习优化

Ahmet H. Güzel, Jenny Seidenschwarz, Benjamin Graham, Jonathan Sadeghi, Jeffrey Hawke, Ilija Bogunovic

机构 * University College London AI Centre(伦敦大学学院人工智能中心) Odyssey University of Basel(巴塞尔大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出一种KL约束的对抗课程,通过训练策略暴露扩散世界模型的高误差轨迹并持续微调,结合优先对抗轨迹缓冲区,解决被动数据中罕见关键转换的鲁棒性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21996 2026-05-29 cs.CV cs.AI 81%

VRAG: Learning World Models for Interactive Video Generation

VRAG:面向交互式视频生成的世界模型学习

Taiye Chen, Xun Hu, Zihan Ding, Chi Jin

机构 * Peking University(北京大学) University of Oxford(牛津大学) Princeton University(普林斯顿大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.CV

AI总结 针对自回归视频生成中累积误差和记忆机制不足的问题,提出视频检索增强生成(VRAG)方法,通过显式全局状态条件降低长期累积误差并提升时空一致性。

Comments Published at NeurIPS 2025. Project page: https://sites.google.com/view/vrag

详情

展开后加载摘要…

URL PDF HTML 收藏