SkeleWAM:用于高效机器人操作的骨架世界-动作建模
SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation
查看机构详情
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
SkeleWAM提出以稀疏3D骨架表示操作场景,结合未来骨架预测,实现高效且参数精简的机器人世界动作建模,在LIBERO-Plus上以57.1M参数达到85.9%成功率。
中文摘要 AI 辅助
世界动作模型(WAMs)将机器人动作生成与未来状态预测相结合。现有的WAMs通常预测视频或学习到的视觉潜在表示,这些表示仅隐式地编码交互几何结构,并且可能保留与控制无关的外观信息。我们提出了SkeleWAM,一种紧凑的WAM,它将操作场景表示为由机器人关节、物体中心和交互点组成的稀疏3D骨架。该骨架从当前的RGB-D观测和机器人本体感觉中在线构建,为动作生成和未来骨架预测提供了统一的几何状态。未来骨架预测为动作学习提供了额外的几何监督,而无需视觉重建。在推理时,SkeleWAM直接从当前骨架和语言指令生成动作,而中值动作共识(MAC)作为随机动作样本的辅助共识策略。在LIBERO-Plus上,SkeleWAM以57.1M参数实现了85.9%的总体成功率,比Cosmos-Policy高出3.7个百分点。这些结果表明,稀疏的3D机器人-物体结构为鲁棒且参数高效的世界动作学习提供了有效的状态空间。项目可在该https URL获取。
英文摘要
World action models (WAMs) combine robot action generation with future state prediction. Existing WAMs typically predict videos or learned visual latents, which represent interaction geometry only implicitly and may retain appearance information unrelated to control. We introduce SkeleWAM, a compact WAM that represents a manipulation scene as a sparse 3D skeleton composed of robot joints, object centers, and interaction points. Constructed online from current RGB-D observations and robot proprioception, the skeleton provides a unified geometric state for action generation and future skeleton prediction. Future skeleton prediction provides additional geometric supervision for action learning without requiring visual reconstruction. At inference, SkeleWAM generates actions directly from the current skeleton and language instruction, while Medoid Action Consensus (MAC) serves as an auxiliary consensus strategy for stochastic action samples. On LIBERO-Plus, SkeleWAM achieves an overall success rate of 85.9% with 57.1M parameters, outperforming Cosmos-Policy by 3.7 percentage points. These results demonstrate that sparse 3D robot--object structure provides an effective state space for robust and parameter-efficient world action learning. The project is available at https://skelewam-project.github.io/.