arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-06-16 至 2026-06-16 共收录 39 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 31 篇

2606.16605 2026-06-16 cs.AI 新提交 94%

ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control

ARB4WM:连续控制中世界模型的对抗鲁棒性基准

Junjian Zhang, Hao Tan, Ruonan Li, Dong Zhu, Aiping Li, Zhaoquan Gu

机构 * College of Computer Science, National University of Defense Technology(国防科技大学计算机学院) College of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) Department of New Networks, Peng Cheng Laboratory(鹏城实验室新型网络部) National Key Laboratory of Advanced Communication Networks(先进通信网络全国重点实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出ARB4WM统一基准,从策略、价值和潜在动力学三个层面评估世界模型在视觉扰动下的对抗鲁棒性,发现多目标攻击和时序暴露模式对安全评估至关重要。

Comments 24 pages, 10 figures, 5 tables. Source code available at https://github.com/zaoanguai/ARB4WM

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16721 2026-06-16 cs.AI 新提交 94%

Medical world models: representing medical states, modelling clinical dynamics and guiding intervention policies

医疗世界模型:表示医疗状态、建模临床动态与指导干预策略

Ke Liu, Mengxuan Li, Yanyi Bao, Tianyun Zhang, Chong Chu, Jiajun Bu, Haishuai Wang

机构 * College of Computer Science, Zhejiang University(浙江大学计算机科学与技术学院) School of Medicine, Zhejiang University(浙江大学医学院) Department of Biomedical Informatics, Harvard University(哈佛大学生物医学信息学系)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出医疗世界模型框架,通过构建患者状态、建模临床动态和支持干预决策,推动医疗AI从静态诊断向动态模拟演进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13053 2026-06-16 cs.RO cs.AI 新提交 94%

EV-WM: Event-Verified World Models for Long-Horizon Robotic Manipulation

EA-WM: 基于任务规范基础的事件感知世界模型用于长时域操作

Kailin Wang, Haoxiang Jie, Yaoyuan Yan, Jiacheng Zhou, Zhiyou Heng

机构 * AI Lab, Country Garden Services Group(碧桂园服务集团AI实验室) Fudan University(复旦大学) Omni AI

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出EA-WM框架,通过事件预测和验证增强预训练特征世界模型,实现长时域操作中任务进展信号的可靠评估与规划。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16489 2026-06-16 cs.LG 新提交 94%

BRICKS-WM: Building Reusability via Interface Composition Kinetics for Structured World Models

BRICKS-WM:通过接口组合动力学构建结构化世界模型的可重用性

Shaowei Zhang, Jiahan Cao, Xunlan Zhou, Shenghua Wan, De-Chuan Zhan

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学计算机软件新技术国家重点实验室) School of Artificial Intelligence, Nanjing University, China(南京大学人工智能学院) School of Intelligence Science and Technology, Nanjing University, China(南京大学智能科学与技术学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出BRICKS-WM框架,将全局动力学分解为通过潜在接口交互的独立模块(如智能体和背景),实现冻结背景模块跨智能体重用,避免从头训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15032 2026-06-16 cs.LG 新提交 94%

How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position

世界模型应如何评估?一个以决策为中心的立场

Yang Yu, Shiyuan Zhang, Yifei Sheng, Haoxiang Ren, Haoxin Lin

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室) School of Artificial Intelligence, Nanjing University(南京大学人工智能学院) Cirquar Technologies

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文指出世界模型评估中声明与证据不匹配的问题,提出以决策为中心的评估框架,强调反事实推理、策略优化等能力,并定义L0-L7评估阶梯。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16286 2026-06-16 cs.LG cs.AI cs.RO 新提交 94%

FlowMPC: Improving Flow Matching policies with World Models

FlowMPC:利用世界模型改进流匹配策略

Chandon Hamel

机构 * Stanford University(斯坦福大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 提出FlowMPC框架,结合流匹配模仿策略与学习的世界模型,通过MPPI规划提升测试时性能,在ManiSkill操作任务中显著提高成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16592 2026-06-16 cs.RO cs.AI cs.CV cs.ET 版本更新 94%

Human Cognition in Machines: A Unified Perspective of World Models

机器中的人类认知:世界模型的统一视角

Timothy Rupprecht, Pu Zhao, Amir Taherin, Arash Akbari, Arman Akbari, Yumei He, Tooba Imtiaz, Sean Duffy, Juyi Lin, Yixiao Chen, Rahul Chowdhury, Enfu Nan, Yixin Shen, Yifan Cao, Haochen Zeng, Weiwei Chen, Geng Yuan, Jennifer Dy, Sarah Ostadabbas, Xuan Zhang, David Kaeli, Edmund Yeh, Yanzhi Wang

机构 * Northeastern University(东北大学) EmbodyX Inc.(EmbodyX公司) Tulane University(路易斯安那州立大学) Cornell University(康奈尔大学) University of Georgia(佐治亚大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出统一框架整合记忆、感知等认知功能,指出动机和元认知研究不足,并引入认知世界模型新类别。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16076 2026-06-16 cs.LG cs.AI cs.GT 新提交 94%

Phys-JEPA: Physics-Informed Latent World Models for Multivariate Time-Series Forecasting

Phys-JEPA:面向多变量时间序列预测的物理信息潜在世界模型

Weizhi Nie, Weichao Liu, Honglin Guo, Yuting Su

机构 * Tianjin University(天津大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出Phys-JEPA架构,将物理一致性约束引入潜在状态和状态转移,分解预测状态为物理和残差分量,在气候、交通、电力数据集上提升预测精度。

Comments Submitted to arXiv as a preliminary manuscript. 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15160 2026-06-16 cs.CV cs.LG 新提交 94%

DLWM: Diverse Latent World Models for Efficient Multimodal Reasoning

DLWM: 多样化潜在世界模型用于高效多模态推理

David Huang, Lianlei Shan

机构 * University of Toronto(多伦多大学) Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出DLWM框架,结合潜在空间推理与强化学习,通过多样化潜在假设和资源感知策略提升多模态推理效率,准确率提升2-5%,内存减少24%。

Comments Preprint. 9 pages main text, 15 pages total including appendix, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16519 2026-06-16 cs.CV 新提交 93%

BadWorld: Adversarial Attacks on World Models

BadWorld:对世界模型的对抗攻击

Linghui Shen, Mingyue Cui, Xingyi Yang

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出BadWorld框架,通过自监督速度攻击和轨迹自适应双层优化,对自回归视觉世界模型进行无标签对抗攻击,暴露其结构脆弱性。

Comments Project Page: https://linghuiishen.github.io/BadWorld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05963 2026-06-16 cs.LG 版本更新 93%

Next-Latent Prediction Transformers Learn Compact World Models

下一潜在预测变换器学习紧凑世界模型

Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward S. Hu, Tim Pearce, Pratyusha Sharma, Akshay Krishnamurthy, Riashat Islam, Alex Lamb, John Langford

机构 * Massachusetts Institute of Technology(麻省理工学院) Microsoft Research(微软研究院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出NextLat方法,通过潜在空间中的自监督预测训练变换器学习紧凑世界模型,提升泛化能力和推理效率。

Comments Microsoft Research Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15594 2026-06-16 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 新提交 93%

Pixels to Proofs: Probabilistically-Safe Latent World Model Control via Parallel Conformal Robust MPC

从像素到证明:通过并行保形鲁棒MPC实现概率安全的潜在世界模型控制

Devesh Nath, Anutam Srinivasan, Haoran Yin, Ruitong Jiang, Jeffrey Fang, Glen Chou

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world-model(abstract)

AI总结 提出SLS^2框架,结合保形预测与鲁棒模型预测控制,在学习的潜在世界模型中实现基于视觉的安全运动规划,提升目标到达性能与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14934 2026-06-16 cs.LG cs.AI 新提交 93%

Separable Neural Architectures as Physical World Models: from Mathematical Theory to Applications

可分离神经架构作为物理世界模型:从数学理论到应用

Reza T Batley, Andrew Kichline, Sourav Saha

机构 * Kevin T. Crofton Department of Aerospace and Ocean Engineering, Virginia Polytechnic Institute and State University(弗吉尼亚理工大学凯文·T·克罗夫顿航空航天与海洋工程系)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 提出可分离神经架构(SNA),结合神经逼近与张量分解,通过变分框架求解偏微分方程,实现高维问题代数级缩放,并在工程案例中取得显著加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19544 2026-06-16 cs.LG cs.RO 版本更新 93%

Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data

通过非策划数据引导世界模型的高效强化学习

Yi Zhao, Aidan Scannell, Wenshuai Zhao, Yuxin Hou, Tianyu Cui, Le Chen, Dieter Büchler, Arno Solin, Juho Kannala, Joni Pajarinen

机构 * Aalto University(阿alto大学) University of Edinburgh(爱丁堡大学) ELLIS Institute Finland(芬兰ELLIS研究所) Deep Render Imperial College London(伦敦帝国理工学院) Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) CIFAR AI Chair(CIFAR人工智能主席) University of Alberta(阿尔伯塔大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所(Amii)) University of Oulu(奥卢大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 提出利用无奖励、混合质量、多本体的非策划离线数据,通过经验回放和执行引导技术解决分布偏移问题,显著提升在线强化学习的样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16576 2026-06-16 cs.CL 新提交 91%

Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning

LLM智能体能否推断世界模型?来自智能体自动机学习的证据

Reef Menaged, Gili Lior, Shauli Ravfogel, Roee Aharoni, Gabriel Stanovsky

机构 * The Hebrew University of Jerusalem(海法大学) New York University(纽约大学) Google Research(谷歌研究)

专题命中 通用世界模型 :world model(title);world models(title);world model(title);world models(title)

AI总结 提出智能体自动机学习框架,通过成员查询和等价查询评估LLM智能体发现隐藏确定性有限自动机的能力,发现性能随DFA规模增加而急剧下降,推理模型优于非推理模型但仍存在规划、整合和假设构建缺陷。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13769 2026-06-16 cs.RO cs.CV cs.LG 新提交 91%

$μ_0$: A Scalable 3D Interaction-Trace World Model

$\mu_0$: 一种可扩展的3D交互轨迹世界模型

Seungjae Lee, Yoonkyo Jung, Jusuk Lee, Jonghun Shin, Amir Hossein Shahidzadeh, Yao-Chih Lee, H. Jin Kim, Jia-Bin Huang, Furong Huang

机构 * University of Maryland, College Park(马里兰大学帕克分校) Seoul National University(首尔大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 提出基于3D轨迹的可扩展世界模型$\mu_0$,通过预测交互点轨迹实现跨本体机器人学习,无需动作标签,性能媲美有监督模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21577 2026-06-16 cs.CL cs.AI cs.LG stat.ML 版本更新 90%

A Unified Definition of Hallucination: It's The World Model, Stupid!

幻觉的统一定义:是世界模型的问题,笨蛋!

Emmy Liu, Varun Gangal, Chelsea Zou, Michael Yu, Xiaoqi Huang, Alex Chang, Zhuofu Tao, Karan Singh, Sachin Kumar, Steven Y. Feng

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出幻觉的统一定义,即用户可观察到的错误内部世界建模,并连接至HalluWorld基准测试,以区分真实幻觉与规划或奖励错误。

Comments ICML 2026. HalluWorld benchmark at https://github.com/DegenAI-Labs/HalluWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16993 2026-06-16 cs.CV 新提交 90%

DreamX-World 1.0: A General-Purpose Interactive World Model

DreamX-World 1.0:通用交互式世界模型

DreamX Team, Yancheng Bai, Rui Chen, Xiangxiang Chu, Rujing Dang, Hao Dou, Bingjie Gao, Qiwen Gu, Siyu Hong, Jiachen Lei, Geng Li, Jifan Li, Ruimin Lin, Qingfeng Shi, Bingze Song, Lei Sun, Jing Tang, Ruitian Tian, Jun Wang, Jiahong Wu, Pengfei Zhang, Shen Zhang, Jiashu Zhu

机构 * DreamX Team(DreamX团队)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);video world model(abstract);video world model(abstract)

AI总结 提出通用交互式文图生视频世界模型DreamX-World 1.0,通过E-PRoPE相机控制、因果强制自回归生成、记忆条件场景持久化和事件指令微调,实现可控长时程生成,在多项指标上超越现有方法。

Comments Project page: https://amap-ml.github.io/DreamX_World, Code: https://github.com/AMAP-ML/DreamX-World

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18428 2026-06-16 cs.RO cs.CV 版本更新 88%

Latent Action Pretraining Through World Modeling

通过世界建模的潜在动作预训练

Bahey Tharwat, Yara Nasser, Ali Abouzeid, Ian Reid

机构 * Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学) Alexandria University(亚历山大大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.CV、cs.RO

AI总结 提出LAWM框架,通过世界建模从无标签视频中学习潜在动作表征,实现跨任务、环境和本体的迁移学习,在LIBERO基准和真实场景中优于使用真实动作预训练的方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12688 2026-06-16 cs.LG cs.AI cs.DC 新提交 87%

M*: A Modular, Extensible, Serving System for Multimodal Models

M*: 一个模块化、可扩展的多模态模型服务系统

Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang

机构 * Stanford University(斯坦福大学) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 提出M*系统,通过将模型表示为数据流图并引入Walk Graph抽象,支持多模态复合模型的高效服务,在多个任务上降低延迟并提升吞吐量。

Comments The codebase is available at https://github.com/mstar-project/mstar

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28185 2026-06-16 cs.CV 版本更新 84%

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

新时代的视觉生成:从原子映射到智能世界建模的演进

Keming Wu, Zuhao Yang, Kaichen Zhang, Shizun Wang, Haowei Zhu, Sicong Leng, Zhongyu Yang, Qijie Wang, Sudong Wang, Ziting Wang, Zili Wang, Hui Zhang, Haonan Wang, Hang Zhou, Yifan Pu, Xingxuan Li, Fangneng Zhan, Bo Li, Lidong Bing, Yuxin Song, Ziwei Liu, Wenhu Chen, Jingdong Wang, Xinchao Wang, Xiaojuan Qi, Shijian Lu, Bin Wang

机构 * Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学) University of Hong Kong(香港大学) National University of Singapore(新加坡国立大学) University of Waterloo(滑铁卢大学) StepFun MiroMind Baidu(百度) Fudan University(复旦大学) Hong Kong University of Science and Technology(香港科技大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) LMMs-Lab(LMMs实验室)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.CV

AI总结 提出五级分类法(原子生成、条件生成、上下文生成、智能生成、世界建模生成),分析流匹配、统一理解与生成模型等关键技术,并指出现有评估高估感知质量而忽视结构、时序和因果缺陷。

Comments Project Page: https://github.com/EvolvingLMMs-Lab/Evolving-Visual-Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18324 2026-06-16 cs.CV cs.AI cs.GR cs.LG stat.ML 版本更新 83%

Improved Baselines with Representation Autoencoders

改进的基于表示自动编码器的基线

Jaskirat Singh, Boyang Zheng, Zongze Wu, Richard Zhang, Eli Shechtman, Saining Xie

机构 * Adobe Research(Adobe研究院) ANU(澳大利亚国立大学) New York University(纽约大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文研究了基于表示自动编码器(RAE)的设计选择,发现三个见解,简化并改进了RAE。首先,研究了一种通用公式,将表示定义为最后k个编码器层的总和,而不是仅最终层。其次,研究了RAE与表示对齐(REPA)的假设,发现两者具有互补的工作机制。最后,改进了RAE在无分类器指导(CFG)中的表现,通过重新参数化DiT模型输出,实现了无需训练第二个模型的指导效果。RAEv2在ImageNet-256上达到了1.06的gFID,且训练效率显著提高。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06212 2026-06-16 cs.CV cs.AI 版本更新 82%

Akasha 2: Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architectur

Akasha 2: 哈密顿状态空间对偶与视觉-语言联合嵌入预测架构

Yani Meziani

机构 * Independent AI Researcher(独立AI研究员) Québec (QC), Canada(魁北克(QC),加拿大)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出 Akasha 2 多模态架构,结合哈密顿状态空间对偶与视觉-语言联合嵌入预测,通过稀疏混合哈密顿专家和哈密顿流匹配实现超低延迟视频预测与合成,在保持能量守恒下取得 SOTA 性能。

Comments No supporting claims were validated in this automated agentic R&D research run

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15162 2026-06-16 cs.CV 新提交 81%

GeoStream: Toward Precise Camera Controlled Streaming Video Generation

GeoStream:迈向精确相机控制的流式视频生成

Yizhou Zhao, Yifan Wang, Xiaoyuan Wang, Yushu Wu, Hao Zhang, Moayed Haji-Ali, Rameen Abdal, Ashkan Mirzaei, Yanyu Li, Willi Menapace, Laszlo Jeni, Sergey Tulyakov, Peter Wonka, Chaoyang Wang

机构 * CMU(卡内基梅隆大学) Northeastern University(东北大学) UIUC(伊利诺伊大学厄巴纳-香槟分校) Rice University(莱斯大学) Snap Inc.(Snap公司) KAUST(阿卜杜拉国王科技大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出GeoStream框架,通过自刷新3D缓存和在线策略蒸馏,实现自回归流式视频生成中的精确度量级相机控制,解决了现有方法在视角移动时控制失效和分布偏移问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13607 2026-06-16 cs.AI 新提交 81%

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning

推理即模式匹配:人类与LLM日常推理中的共享机制

Zach Studdiford, Gary Lupyan

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 研究通过比较人类和25个LLM在日常因果推理中的错误模式,发现两者均表现出模式匹配而非抽象世界模型驱动的推理,并识别出LLM中驱动响应的注意力头可预测人类推理错误。

Comments 13 pages main text, 51 pages supplementary text

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14879 2026-06-16 cs.RO cs.CV cs.LG 新提交 73%

VANDERER: Map-Free Exploration using Future-Aware and Visual-Curiosity-Guided Diffusion Policy

VANDERER: 基于未来感知与视觉好奇心引导扩散策略的无地图探索

Venkata Naren Devarakonda, Raktim Gautam Goswami, Prashanth Krishnamurthy, Farshad Khorrami

机构 * Control/Robotics Research Laboratory (CRRL), Department of Electrical and Computer Engineering, NYU Tandon School of Engineering(纽约大学坦登工程学院电气与计算机工程系控制/机器人研究实验室(CRRL)) New York University Abu Dhabi (NYUAD) Center for Artificial Intelligence and Robotics (CAIR)(纽约大学阿布扎比分校人工智能与机器人中心(CAIR))

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG、cs.CV、cs.RO

AI总结 提出VANDERER框架,利用视觉好奇心模块引导预训练扩散策略,仅依赖单目图像实现高效无地图探索,在多种模拟环境中平均探索面积比NoMaD多13.4%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04184 2026-06-16 cs.CV 版本更新 69%

GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs

GroupToM-Bench: 多模态大语言模型中群体心智理论和非线性社会涌现的基准测试

Weidong Tang, Jierui Li, Yueling Hou, Zihan Mei, Can Zhang, Xinyan Wan, Zhiyuan Liang, Pengfei Zhou, Yang You, Wangbo Zhao

机构 * Xidian University(西安电子科技大学) National University of Singapore(新加坡国立大学) University of Electronic Science and Technology of China(电子科技大学) University of Science and Technology of China(中国科学技术大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 针对多模态大语言模型在群体心智理论推理上的不足,提出GroupToM-Bench基准,通过七级认知审计框架评估模型从微观BDI状态到宏观结果预测的因果链,揭示模型在处理社会结构和非线性集体动态上的缺陷。

Comments ACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23234 2026-06-16 cs.LG cs.CY 版本更新 67%

Assessing Predictive Models for Fairness Based on Movement Patterns

基于移动模式评估预测模型的公平性

Francesco Lettich, Mario A. Nascimento, Chiara Pugliese, Chiara Renso

机构 * University of Padua(帕多瓦大学)

专题命中 通用世界模型 :predictive model(title,abstract);predictive models(title,abstract);分类 cs.LG

AI总结 针对预测模型的空间公平性,提出将公平性概念从单一地理位置扩展到移动模式,并采用空间扫描统计方法检测基于移动模式的不公平性。

Comments 33 pages, 10 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16570 2026-06-16 cs.RO 新提交 50%

Automated Digital Twin Construction for Highway Scenarios Using LiDAR Point Clouds and OpenStreetMap

基于LiDAR点云和OpenStreetMap的高速公路场景自动数字孪生构建

Yongqi Zhao, Dong Bi, Paul Kovacevic, Tomislav Mihalj, Martin Schabauer, Johannes Betz, Arno Eichberger

机构 * Institute of Automotive Engineering, Graz University of Technology(格拉茨技术大学汽车工程研究所) School of Intelligent Connected Vehicle, Hubei University of Automotive Technology(湖北汽车工业学院智能网联汽车学院) Professorship of Autonomous Vehicle Systems, Technical University of Munich(慕尼黑工业大学自动驾驶系统教席)

专题命中 通用世界模型 :environment model(abstract);分类 cs.RO

AI总结 提出融合LiDAR点云与OpenStreetMap数据的自动化流程,生成地理参考的ASAM OpenDRIVE高速公路地图,实现车道级几何与拓扑的完整建模,平均横向RMSE为0.740米。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15288 2026-06-16 cs.LG cs.AI physics.ao-ph 新提交 50%

Hybrid NARX-LLM for Greenland Iceberg Discharge: Prompt-Driven Residual Correction

混合NARX-LLM用于格陵兰冰山排放:提示驱动的残差校正

Yiquan Gao, Duohui Xu

机构 * Heriot-Watt University(赫瑞瓦特大学) StudioYG

专题命中 通用世界模型 :分类 cs.AI、cs.LG;predictive model(abstract);predictive models(abstract)

AI总结 提出混合NARX-LLM框架,结合非线性自回归模型与大型语言模型进行残差校正,并引入物理信息提示方法,用于建模格陵兰冰山排放的复杂非线性动态,提升预测准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏