世界模型能造就更好的机器人吗?预测性具身智能评估基准综述
Do World Models Make Better Robots? A Survey of Evaluation Benchmarks for Predictive Embodied Intelligence
浏览论文内容
中文总结 AI 辅助
针对世界模型是否带来闭环优势的问题,本综述整理了160个基准,指出测量缺口,并提出隔离预测优势的评估循环和四个优势感知指标。
中文摘要 AI 辅助
机器人学习现在沿着两条很少交汇的轨道发展。一方面,直接的视觉-语言-动作(VLA)策略将观测映射到动作,并通过闭环任务成功率进行评分。另一方面,预测性和生成式世界模型预测未来的观测,并通过开环预测或生成质量进行评分。一个自然的问题介于两者之间:世界建模是否比直接策略带来可测量的闭环优势,并且针对哪些机器人能力?我们认为,该领域目前无法回答这个问题,原因在于测量方式上的差距,而非模型本身。世界模型基准在评分预测时从未执行预测,而任务成功率套件只承载单一策略,从未构建世界模型与VLA的对比。本综述围绕这一差距描绘评估格局。我们整理了2017年至2026年间160个经网络验证的基准,并按评估模式、机器人能力和模型系列将其组织为四条通道:策略套件、具身智能体、世界模型评估和预测到动作的桥梁。在整个语料库中,160个基准中有138个与模型无关,只有11个(7%)构建了明确的VLA与世界模型对比;反事实能力几乎完全未被测量,只有四个基准将预测转化为执行的动作。我们贡献了一个操作性分类法、与最近的八篇综述的覆盖范围比较(我们的综述是唯一将能力与模型系列交叉的)、一个隔离预测优势的评估循环,以及一个包含四个优势感知指标的可操作协议,这些指标锚定在指定的测试平台上。核心主张不是世界模型有帮助或没有帮助,而是回答这个问题需要构建旨在提出该问题的基准。
英文摘要
Robot learning now advances along two tracks that rarely meet. On one side, direct Vision-Language-Action (VLA) policies map observations to actions and are scored by closed-loop task success. On the other, predictive and generative world models forecast future observations and are scored by open-loop prediction or generation quality. A natural question sits between them: does world modelling earn a measurable, closed-loop advantage over a direct policy, and for which robotic capabilities? We argue that the field cannot yet answer this question, and that the reason is a gap in how it is measured, not in the models themselves. World-model benchmarks score prediction without ever executing it, while task-success suites host a single policy and never build a world-model versus VLA contrast. This survey maps the evaluation landscape around that gap. We catalogue 160 web-verified benchmarks spanning 2017 to 2026 and organise them by evaluation mode, robotic capability, and model family into four lanes: policy suites, embodied agents, world model evaluation, and prediction-to-action bridges. Across the corpus, 138 of 160 benchmarks are model-agnostic and only 11 (7%) build an explicit VLA-versus-world-model contrast; counterfactual capability is almost entirely unmeasured, and only four benchmarks turn prediction into executed action. We contribute an operational taxonomy, a coverage comparison against the eight closest surveys (ours is the only one to cross capability with model family), an evaluation loop that isolates the advantage of prediction, and an actionable protocol of four advantage-aware metrics anchored on named testbeds. The organising claim is not that world models help or do not help, but that answering the question requires benchmarks built to ask it.
发表机构
- UC Berkeley(加州大学伯克利分校)
- San Jose State University(圣何塞州立大学)
- Meta
- Apple(苹果公司)
- PocketFM
- Pragya Lab, BITS Pilani Goa(BITS Pilani Goa 的 Pragya 实验室)
机构由 AI 辅助整理,请以论文原文为准。