具身智能的世界模型:从合理到可控再到可行动
World Models for Embodied Intelligence: From Plausible to Controllable to Actionable
浏览论文内容
中文总结 AI 辅助
本文提出具身智能世界模型的三级能力框架(合理、可控、可行动),并配以3×4矩阵,系统综述各领域进展,推动评估从视觉保真转向行为改善。
中文摘要 AI 辅助
世界模型通过维持隐藏状态、预测后果、比较干预措施以及在执行偏离预期时进行调整,将具身智能中的感知与决策联系起来。尽管进展通常以视觉保真度来衡量,但其价值在于改善行为。在伸手拿杯子之前,人会预判其重量和抓取阻力,从而在接触前调整手型。这种预判是粗略的,很少是图像式的,却能指导行动。这引出了一个核心问题:哪些预测能力能改善行为?现有的综述按架构、输出模态或应用领域组织,使这一问题隐而不显。我们引入了三个逐步增强的能力层级:合理模型保留与任务相关的时间、几何或物理结构;可控模型额外预测干预如何改变该结构;可行动模型将预测转化为规划、行动、学习、评估、验证、恢复或数据选择中的可衡量收益。我们用3×4矩阵补充这一层级,该矩阵将几何、物理和行动锚定与围绕数据、奖励、策略和模型本身的改进循环交叉。利用这一框架,我们综述了操作、导航、 locomotion、自动驾驶和通用具身学习,追溯技术进展,阐明能力要求,并审视数据集、基准和评估协议。我们识别了长时程一致性、不确定性校准、因果干预测试、延迟、验证与恢复以及跨具身迁移方面的挑战。这一视角将评估从视觉合理性转向预测是否捕获任务相关状态、反映干预效果并改善具身代理的闭环行为。
英文摘要
World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interventions, and adapting when execution departs from expectations. Although progress is often measured by visual fidelity, their value lies in improving behavior. Before reaching for a cup, a person anticipates its weight and resistance to grasping, shaping the hand before contact. Such anticipation is coarse and rarely pictorial, yet it guides action. This raises a central question: which predictive capabilities improve behavior? Existing surveys, organized by architecture, output modality, or application domain, leave this question implicit. We introduce three progressively stronger capability levels: Plausible models preserve task-relevant temporal, geometric, or physical structure; Controllable models additionally predict how interventions alter that structure; and Actionable models translate predictions into measurable gains in planning, action, learning, evaluation, verification, recovery, or data selection. We complement this hierarchy with a 3 x 4 matrix crossing geometry, physics, and action grounding with improvement loops centered on data, rewards, policies, and the model itself. Using this framework, we survey manipulation, navigation, locomotion, autonomous driving, and general embodied learning, tracing technical progressions, clarifying capability requirements, and examining datasets, benchmarks, and evaluation protocols. We identify challenges in long-horizon consistency, uncertainty calibration, causal intervention testing, latency, verification and recovery, and cross-embodiment transfer. This perspective shifts evaluation from visual plausibility toward whether predictions capture task-relevant state, reflect intervention effects, and improve the closed-loop behavior of embodied agents.
发表机构
- HKUST(GZ)(香港科技大学(广州))
- University of Oxford(牛津大学)
- Princeton University(普林斯顿大学)
- Tencent(腾讯)
- Sun Yat-sen University(中山大学)
- Shanghai Jiao Tong University(上海交通大学)
- Peking University(北京大学)
- Tsinghua University(清华大学)
- Alibaba Group(阿里巴巴集团)
- Singapore Management University(新加坡管理大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。