XEWorld:基于动作的世界模型能否泛化至未见过的机器人 embodiment?
XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?
浏览论文内容
中文总结 AI 辅助
该研究针对XEWorld测试平台,发现基于动作的世界模型因视觉与物理动力学解耦不足,难以跨机器人 embodiment 泛化,需架构创新突破瓶颈。
中文摘要 AI 辅助
基于动作的世界模型是用于机器人操纵的有前景的学习模拟器,但仅在训练机器人上评估它们无法揭示其是否捕捉物理动力学,还是仅记忆视觉模式。为回答模型能否忠实地渲染从未见过的机器人,我们引入XEWorld,这是一个受控的跨 embodiment 世界模型测试平台,通过在物理完全相同的场景中评估保留的机器人来分离不同 embodiment。我们的系统分析揭示了一个共同的架构瓶颈:当前模型主要充当2D视觉模式匹配器,其泛化由视觉相似性而非运动学相似性决定。受此限制,它们难以将抽象的数值关节动作转换为连贯的视觉轨迹,也无法从静态初始观测预测动态视觉变化。因此,成功零样本渲染未见过的 embodiment 严格需要大量基于物理的线索,特别是像素空间动作和显式时空对齐。即使通过少样本适应绕过零样本障碍,强制外观恢复也会触发对已见过 embodiment 的灾难性遗忘。这些失败共同暴露了将学习到的物理动力学应用于新视觉外观的关键无能,强调实现真正的跨 embodiment 泛化需要将视觉外观与底层物理动力学解耦的架构创新。
英文摘要
Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual patterns. To answer whether a model can faithfully render a robot it has never seen, we introduce XEWorld, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes. Our systematic analysis uncovers a shared architectural bottleneck: current models act primarily as 2D visual pattern matchers whose generalization is governed by visual similarity rather than physical kinematic similarity. Driven by this limitation, they struggle to translate abstract numeric joint actions into coherent visual trajectories, and fail to predict dynamic visual changes from static initial observations. Consequently, successfully rendering an unseen embodiment zero-shot strictly requires heavily grounded cues, specifically pixel-space actions and explicit spatial-temporal alignment. Even when bypassing this zero-shot barrier via few-shot adaptation, the forced appearance recovery triggers catastrophic forgetting of seen embodiments. Together, these failures expose a critical inability to apply learned physical dynamics to novel visual appearances, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying physical dynamics.