arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-05-12 至 2026-05-12 共收录 27 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 20 篇

2605.10858 2026-05-12 cs.CV cs.RO 95%

Is Your Driving World Model an All-Around Player?

你的驾驶世界模型是全能选手吗?

Lingdong Kong, Ao Liang, Tianyi Yan, Hongsi Liu, Wesley Yang, Ziqi Huang, Xian Sun, Wei Yin, Jialong Zuo, Yixuan Hu, Dekai Zhu, Dongyue Lu, Youquan Liu, Guangfeng Jiang, Linfeng Li, Xiangtai Li, Long Zhuo, Lai Xing Ng, Benoit R. Cottereau, Changxin Gao, Liang Pan, Wei Tsang Ooi, Ziwei Liu

专题命中 通用世界模型 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 本文提出WorldLens基准测试,评估驾驶世界模型在视觉和行为真实性方面的综合表现,揭示现有模型在不同维度上的不足,并引入WorldLens-26K和WorldLens-Agent提升评估的可解释性。

Comments CVPR 2026 VideoWorldModel Workshop; Project Page at https://worldbench.github.io/worldlens GitHub at https://github.com/worldbench/WorldLens

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08279 2026-05-12 cs.LG cs.AI 95%

LaWM: Least Action World Models for Long-Horizon Physical Consistency from Visual Observations

LaWM:基于视觉观测的最短作用世界模型用于长时间物理一致性

Qixin Xiao, Maani Ghaffari

机构 * University of Michigan(密歇根大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 LaWM通过在视觉潜在空间中实现最小作用原理,提升长horizon视觉预测的物理一致性,改进了物理不变性、背景一致性、运动平滑度和外观几何预测指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10806 2026-05-12 cs.CV cs.AI cs.LG 94%

PhyGround: Benchmarking Physical Reasoning in Generative World Models

PhyGround:用于生成世界模型中物理推理的基准测试

Juyi Lin, Arash Akbari, Yumei He, Lin Zhao, Haichao Zhang, Arman Akbari, Xingchen Xu, Zoe Y. Lu, Enfu Nan, Hokin Deng, Edmund Yeh, Sarah Ostadabbas, Yun Fu, Jennifer Dy, Pu Zhao, Yanzhi Wang

机构 * Northeastern University(东北大学) Tulane University(路易斯安那州立大学) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 PhyGround通过250个精心设计的提示和13种物理定律的分类,评估视频生成中的物理推理能力,采用大规模人工研究验证并发布专用VLM评估工具PhyJudge-9B,显著降低偏见。

Comments Preprint. 56 pages, 39 figures, 40 tables. Project page: https://phyground.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09241 2026-05-12 cs.LG cs.AI 94%

Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models

Sub-JEPA:子空间高斯正则化用于稳定的端到端世界模型

Kai Zhao, Dongliang Nie, Yuchen Lin, Zhehan Luo, Yixiao Gu, Deng-Ping Fan, Dan Zeng

机构 * Shanghai University(上海大学) The University of Manchester(曼彻斯特大学) Nankai University(南开大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 Sub-JEPA通过在多个随机子空间中应用高斯约束,缓解了JEPA训练中的偏差-方差权衡问题,提升了训练稳定性与表示灵活性,实验显示其在连续控制环境中优于LeWM。

Comments https://github.com/intcomp/Sub-JEPA

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08578 2026-05-12 cs.LG cs.AI 94%

Probing the Impact of Scale on Data-Efficient, Generalist Transformer World Models for Atari

探测数据高效、通用型Transformer世界模型在Atari上的尺度影响

Jooyeon Kim

机构 * Graduate School of Artificial Intelligence(人工智能研究生院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文通过最小化Transformer世界模型分析Atari 100k基准的缩放行为,发现环境存在不同的缩放区域,联合训练可稳定缩放动态并提升下游控制性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08954 2026-05-12 cs.LG cs.AI 93%

MolWorld: Molecule World Models for Actionable Molecular Optimization

MolWorld: 用于可操作分子优化的分子世界模型

Yang Qiao, Bo Pan, Hao-Wei Pang, Peter Zhiping Zhang, Liying Zhang, Liang Zhao

机构 * Emory University(埃默里大学) Merck & Co., Inc.(默克公司)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 MolWorld通过分子转移图的序列扩展实现可操作分子优化,通过学习世界模型增强分子结构连接性,提升分子设计的可行性和连续性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09886 2026-05-12 cs.RO 92%

Network-Efficient World Model Token Streaming

网络高效的世界模型令牌流

Shatadal Mishra, Ahmadreza Moradipari, Nejib Ammar

机构 * InfoTech Labs, Toyota Motor North America R\&D, Mountain View, CA, USA

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);driving world model(abstract)

AI总结 本文研究了在分布式计算和连接车辆中高效传输离散世界模型状态的方法,提出了一种自适应算法,通过余弦距离优先更新令牌并触发关键帧,从而在带宽受限环境下提升同步效率和下游令牌动态的实用性。

Comments Accepted at IEEE VNC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11537 2026-05-12 cs.LG cs.AI 90%

Simulus: Combining Improvements in Sample-Efficient World Model Agents

Simulus:结合高效样本世界模型代理的改进

Lior Cohen, Kaixin Wang, Bingyi Kang, Uri Gadot, Shie Mannor

机构 * Technion(技术ion大学) Microsoft Research(微软研究院) ByteDance Seed(字节跳动种子实验室)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出Simulus,一种模块化的基于标记的世界模型代理,通过整合灵活的标记化框架、内在动机、优先世界模型回放和回归-分类方法,提升了样本效率,适用于多个基准测试。

Comments Revised version: updated title, abstract, and framing to better reflect our contributions and situate the work within the literature

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09131 2026-05-12 cs.AI cs.MA 90%

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

MCP-Cosmos:用于MCP环境复杂任务执行的世界模型增强型智能体

Giridhar Ganapavarapu, Dhaval Patel

机构 * IBM

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 MCP-Cosmos融合世界模型与MCP框架,通过引入生成世界模型提升智能体在复杂任务中的执行能力,实验显示其在工具成功率和参数准确性等方面有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08268 2026-05-12 cs.MA cs.AI 87%

Insider Attacks in Multi-Agent LLM Consensus Systems

多智能体大语言模型共识系统中的内部攻击

Xiaolin Sun, Zixuan Liu, Yibin Hu, Zizhan Zheng

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,Location,Country) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,Location,Country) Department of Computer Science, Tulane University, New Orleans, United States of America(计算机科学系, Tulane大学,新奥尔良,美国)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 研究多智能体大语言模型共识系统中的内部攻击问题,提出基于世界模型的框架,通过学习良性智能体的潜在行为状态并利用强化学习训练攻击者,有效降低共识率并延长分歧时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10404 2026-05-12 cs.CV 81%

Position: Life-Logging Video Streams Make the Privacy-Utility Trade-off Inevitable

位置:生命记录视频流使隐私与效用的权衡不可避免

Tianyuan Zou, Liang Yue, Yang Liu, Ya-Qin Zhang, Sijie Cheng

机构 * Institute for AI Industry Research, Tsinghua University, Beijing, China(清华大学人工智能产业研究院) RayNeo.AI, Shenzhen, China(深圳RayNeo.AI) Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 生命记录视频流虽提升AI系统效用,但暴露敏感信息引发隐私风险,现有保护措施无法兼顾效用与隐私,需进一步研究权衡设计与标准化评估。

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08956 2026-05-12 cs.AI 81%

Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

代理式AI科学家并非为自主科学发现而建

Harshit Bisht, Vinay Kumar, Kevin Maik Jablonka, Mausam, N. M. Anoop Krishnan

机构 * Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(印度理工学院德里人工智能学院) Department of Computer Science and Engineering, Indian Institute of Technology Delhi(印度理工学院德里计算机科学与工程系) Department of Civil and Environmental Engineering, Indian Institute of Technology Delhi(印度理工学院德里土木与环境工程系) Laboratory of Organic and Macromolecular Chemistry (IOMC), Friedrich Schiller University Jena(耶拿弗里德里希·席勒大学有机与大分子化学实验室) Center for Energy and Environmental Chemistry Jena, Friedrich Schiller University Jena(耶拿弗里德里希·席勒大学能源与环境化学中心) Helmholtz Institute for Polymers in Energy Applications Jena (HIPOLE Jena)(耶拿海德堡聚合物能源应用研究所(HIPOLE耶拿))

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文指出,尽管代理式AI科学家能作为合作者,但其设计并非为自主科学发现。挑战包括问题选择偏差、实验室知识缺失、输出多样性压缩及缺乏实验反馈,需重新审视基础设计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23629 2026-05-12 cs.GR 80%

From Visual Synthesis to Interactive Worlds: Toward Production-Ready 3D Asset Generation

从视觉合成到交互世界:迈向可生产3D资产生成

Jiafeng Wu, Zhuofan Lou, Jian Liu, Dazhao Du, Chunchao Guo, Song Guo

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文探讨了从孤立视觉形状到可部署的结构化资产的3D内容生成进展,分析了游戏开发、AI、世界模拟等对3D资产的需求,提出基于资产生产流程的分类方法,评估现有方法的生成质量和可用性,并指出数据质量、生成可控性等挑战。

Comments Preprint. Jiafeng Wu and Zhuofan Lou contributed equally. Project page: https://christinebobby.github.io/production-ready-3d-survey/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09693 2026-05-12 cs.CV cs.AI cs.LG 73%

Do multimodal models imagine electric sheep?

多模态模型是否想象出电羊?

Santhosh Kumar Ramakrishnan, Carl Vondrick, Raja Giryes, Philipp Krähenbühl, Vladlen Koltun

机构 * Apple(苹果公司)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG、cs.CV

AI总结 研究发现多模态模型在解决空间谜题时会形成心理图像,通过微调Qwen3.5 VLM解决多种视觉推理任务,发现动作序列预测能提升解谜准确率,尤其在需要空间推理的任务中效果显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09790 2026-05-12 cs.DC cs.AI cs.LG 71%

Multi-Tier Labeling and Physics-Informed Learning for Orbital Anomaly Detection at Scale

多层级标签与物理引导学习用于大规模轨道异常检测

Yong Fu

机构 * Substratum Labs, Inc(Substring实验室)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出多层级标签方法,结合物理规则、IMM-UKF和补充元素校准,实现大规模轨道异常检测,通过Transformer模型提升召回率,探讨神经ODE在轨道世界模型中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09650 2026-05-12 cs.AI cs.LG 71%

Workspace Optimization: How to Train Your Agent

工作区优化:如何训练你的智能体

Elad Sarafian, Gal Kaplun, Ron Banner, Daniel Soudry, Boris Ginsburg

机构 * NVIDIA

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过优化智能体的工作区结构来提升其在多轮任务中的表现,通过模拟权重空间训练方法,利用人工制品、证据、反例和文本反馈来改进智能体的执行能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09678 2026-05-12 cs.AI 69%

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities

荒诞世界:一种简单却强大的方法,用于将现实世界扭曲以探测LLM推理能力

Ryan Albright, Golam Md Muktadir, Zarif Ikram, S M Jubaer, Mehrab Hossain, Dianbo Liu

机构 * The Nueva School(新维学校) University of Southern California(南加州大学) Notre Dame College(诺特大学) Arizona State University(亚利桑那州立大学) National University of Singapore(新加坡国立大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 本文提出Absurd World框架,通过扭曲现实世界来测试LLM的推理能力,验证其在简单逻辑任务中的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09355 2026-05-12 cs.LG 69%

FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

FLAME:适应性专家混合用于连续多模态多任务学习

Xing Han, Shravan Chaudhari, Tanvi Ranade, Rama Chellappa, Suchi Saria

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 FLAME提出了一种可扩展的专家混合框架,支持多任务预训练和连续学习,通过模态特定路由器处理不同模态的任务,同时利用低秩内存子空间压缩专家知识,提升参数效率和减少灾难性遗忘。

Comments 37 pages, 25 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.04899 2026-05-12 cs.LG 69%

A geometric relation of the error introduced by sampling a language model's output distribution to its internal state

语言模型输出分布采样引入的误差与内部状态的几何关系

Albert F. Modenbach

机构 * Department of Mathematics, King's College London, United Kingdom(伦敦国王学院数学系)

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.LG

AI总结 研究揭示语言模型输出分布采样误差与内部状态几何结构的关系,通过几何方法分析token嵌入空间,发现其曲率在棋类任务中反映模型对问题的内部表示。

Comments 12 Pages, 10 Figures, 2 Appendices. To appear in Proceedings of ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08958 2026-05-12 cs.LG 60%

Learning predictive models for combinations of heterogeneous proteomic data sources

学习异质蛋白质组数据源的预测模型

Michal Valko, Richard Pelikan, Miloš Hauskrecht

专题命中 通用世界模型 :predictive model(title);predictive models(title);分类 cs.LG

AI总结 本文研究了两种蛋白质混合物表达数据源的个体和联合效用,提出了一种模型融合方法以最大化异质数据集的结合效益。

Comments Published at in AMIA Summit on Translational Bioinformatics (STB 2008

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 自动驾驶 2 篇

2605.09701 2026-05-12 cs.CV 93%

DriveFuture: Future-Aware Latent World Models for Autonomous Driving

DriveFuture: 用于自动驾驶的面向未来的潜在世界模型

Yufeng Hong, Xiaotian Zhou, Yingyan Li, Xiangpo Zhou, Lin Liu, Yadan Luo, Shaoqing Xu, Lei Yang, Ziying Song

机构 * Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beihang University(北航) Beijing Jiaotong University(北京交通大学) The University of Queensland(昆士兰大学) University of Macau(澳门大学) Nanyang Technological University(南洋理工大学) School of Artificial Intelligence ( School of Software), Yanshan University(燕山大学人工智能学院(软件学院))

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出DriveFuture,一种面向未来的潜在世界建模框架,通过将未来世界状态条件化于当前潜在状态建模过程,提升自动驾驶轨迹规划性能。

Comments 24pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10564 2026-05-12 cs.CV cs.RO 90%

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving

DeepSight: 通过潜在状态预测实现长视距世界建模的端到端自动驾驶

Lingjun Zhang, Changjie Wu, Linzhe Shi, Jiangyang Li, Jiaxin Liu, Lei Yang, Hang Zhang, Mu Xu, Hong Wang

机构 * Tsinghua University(清华大学) Amap, Alibaba Group(阿里巴巴集团Amap) Nanyang Technological University(南洋理工大学)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);driving world model(abstract);driving world model(abstract)

AI总结 本文提出通过鸟瞰图空间预测连续未来帧的潜在语义特征,实现长视距世界建模,并引入高效适应性文本推理机制提升复杂场景下的驾驶性能。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 模型式强化学习 3 篇

2311.03600 2026-05-12 cs.RO 53%

Scalable and Efficient Continual Learning from Demonstration via a Hypernetwork-generated Stable Dynamics Model

通过超网络生成的稳定动力学模型实现可扩展且高效的示范学习

Sayantan Auddy, Jakob Hollenstein, Matteo Saveriano, Antonio Rodríguez-Sánchez, Justus Piater

机构 * Faculty of Electrical Engineering and Computer Science, Technical University of Berlin(电气工程与计算机科学系,柏林技术大学) Department of Computer Science, University of Innsbruck(计算机科学系,因斯布鲁克大学) Digital Science Center (DiSC), University of Innsbruck(数字科学中心(DiSC),因斯布鲁克大学) Department of Industrial Engineering, University of Trento(工业工程系,特伦托大学) Singular Research Center on Intelligent Systems (CiTIUS), University of Santiago de Compostela(智能系统研究中心(CiTIUS),圣地亚哥-德孔波斯特拉大学)

专题命中 模型式强化学习 :dynamics model(title,abstract);分类 cs.RO

AI总结 本文提出一种稳定的持续学习方法,利用超网络生成轨迹学习动力学模型和Lyapunov函数,提升稳定性与准确性,通过任务嵌入减少训练时间,实验证明其在多任务学习中优于现有方法。

Comments To appear in IEEE Transactions on Cognitive and Developmental Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04233 2026-05-12 cs.LG cs.AI 53%

PAINET: A Principled Efficient Transformer for 3D Dynamics Modeling

PAINET:一种基于原理的高效变压器用于3D动态建模

Kai Yang, Yuqi Huang, Junheng Tao, Wanyu Wang, Qitian Wu

机构 * Department of Computer Science and Engineering, Shanghai Jiao Tong University(上海交通大学计算机科学与工程系) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) SJTU Paris Elite Institute of Technology, Shanghai Jiao Tong University(上海交通大学巴黎精英理工学院) Department of Physics, Harvard University(哈佛大学物理系) Harvard-MIT Center for Ultracold Atoms(哈佛-麻省理工冷原子中心) Eric and Wendy Schmidt Center, Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所埃里克和wendy Schmidt中心)

专题命中 模型式强化学习 :dynamics model(title);分类 cs.AI、cs.LG

AI总结 本文提出PAINET,一种基于SE(3)等价性的变压器,用于学习多体系统的所有对相互作用,通过物理启发的注意力网络和并行解码器,在多个真实世界基准上实现了3D动态预测的误差降低。

Comments 24 pages, published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16741 2026-05-12 cs.LG math.OC stat.ML 50%

Meta-reinforcement learning with minimum attention

最小注意力的元强化学习

Shashank Gupta, Pilhwa Lee

机构 * Department of EECS(电气与计算机工程系) Department of Mathematics(数学系) University of Michigan(密歇根大学) Morgan State University(摩根州立大学)

专题命中 模型式强化学习 :model-based RL(abstract);分类 cs.LG

AI总结 本文研究了将最小注意力应用于强化学习作为奖励的一部分,探索其与元学习和稳定性的联系,通过模型基于元学习和梯度基元策略学习,在高维非线性动态中实现高效适应与能量效率提升。

Comments 30 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 仿真与规划 2 篇

2604.18486 2026-05-12 cs.CV cs.CL cs.RO 71%

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

小米OneVL:基于视觉-语言解释的一步潜在推理与规划

Jinghui Lu, Jiayi Guan, Zhijian Huang, Jinlong Li, Guang Li, Lingdong Kong, Yingyan Li, Han Wang, Shaoqing Xu, Yuechen Luo, Fang Li, Chenxu Dang, Junli Wang, Tao Xu, Jing Wu, Jianhua Wu, Xiaoshuai Hao, Wen Zhang, Tianyi Jiang, Lingfeng Zhang, Lei Zhou, Yingbo Tang, Jie Wang, Yinfeng Gao, Xizhou Bu, Haochen Tian, Yihang Qiu, Feiyang Jia, Lin Liu, Yigu Ge, Hanbing Li, Yuannan Shen, Jianwei Cui, Hongwei Xie, Bing Wang, Haiyang Sun, Jingwei Zhao, Jiahui Huang, Pei Liu, Zeyu Zhu, Yuncheng Jiang, Zibin Guo, Chuhong Gong, Hanchao Leng, Kun Ma, Naiyan Wang, Guang Chen, Kuiyuan Yang, Hangjun Ye, Long Chen

机构 * Xiaomi Embodied Intelligence Team(小米具身智能团队)

专题命中 仿真与规划 :world model(abstract);world model(abstract);分类 cs.CV、cs.RO

AI总结 OneVL通过双重辅助解码器监督的紧凑潜在标记,实现了视觉-语言解释的一步潜在推理与规划,首次在延迟条件下超越了显式推理方法。

Comments Technical Report; 49 pages, 22 figures, 10 tables; Project Page at https://xiaomi-embodied-intelligence.github.io/OneVL GitHub at https://github.com/xiaomi-research/onevl

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10166 2026-05-12 cs.RO 69%

Data-Asymmetric Latent Imagination and Reranking for 3D Robotic Imitation Learning

数据不对称潜在想象与重排序在3D机器人模仿学习中的应用

Lianghao Luo, Xizhou Bu, Ruyan Liu, Qingqiu Huang, Chufeng Tang, Xiaoshuai Hao, Hongbo Wang, Wei Li

专题命中 仿真与规划 :world model(abstract);world model(abstract);分类 cs.RO

AI总结 本文提出DALI-R框架,通过潜在世界模型和任务完成评分器,利用混合质量轨迹提升3D机器人模仿学习的决策能力,无需额外高质量示范。

详情

展开后加载摘要…

URL PDF HTML 收藏