Copper-Policy:聚焦表示以实现稳健的机器人操作
Copper-Policy: Focus on the Representation for Robust Robot Manipulation
浏览论文内容
中文总结 AI 辅助
Copper-Policy通过联合嵌入预测学习紧凑世界表示,无需像素重建,实现高效训练(2B参数9.67小时),在模拟和真实机器人任务中达到或超越现有方法,兼顾性能与效率。
中文摘要 AI 辅助
世界动作模型(WAMs)通过建模未来场景演化来获取行为先验,但在像素或潜在空间中预测详细的未来会带来巨大成本。最近的证据表明,在没有测试时生成的情况下,联合训练的优势仍然存在,这引发了一个问题:WAM必须学习什么才能改进控制?我们提出了Copper-Policy,它学习一个紧凑的世界表示,并与策略一起使用,而不是依赖于预定义的目标空间。通过时间联合嵌入预测,它基于任务意图预测未来的观测嵌入,而无需重建像素。这种预测和动作解码共同塑造了表示,而策略保留了对当前帧空间细节的访问以执行动作。表示分析表明,学习到的特征能更好地区分任务驱动的变化与扰动,并为控制提供互补信息。紧凑的预测目标减少了每个样本的训练令牌数量,使得一个2B参数的模型在8块RTX 5090 GPU上仅用9.67小时完成训练,并且在匹配的A100 GPU上比Fast-WAM快6倍。在RoboTwin上,Copper-Policy在无具身预训练的情况下优于所有对比方法,并在LIBERO-Plus(80.85%)上优于几种具身预训练的VLA。在三个具有挑战性的真实机器人任务中,它的表现与π0.5相当,并获得了更高的平均分数。这些结果共同表明,Copper-Policy结合了强大的控制性能与高效的训练。
英文摘要
World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a question: what must a WAM learn to improve control? We introduce Copper-Policy, which learns a compact World representation with the policy rather than relying on a predefined target space. Through temporal joint-embedding prediction, it predicts future observation embeddings conditioned on task intention without reconstructing pixels. This prediction and action decoding shape the representation jointly, while the policy retains access to current-frame spatial detail for execution. Representation analyses show that the learned features better separate task-driven change from perturbations and provide complementary information for control. Compact prediction targets reduce training tokens per sample, enabling a 2B-parameter model trained in 9.67 hours on 8$\times$ RTX 5090 GPUs and 6$\times$ faster than Fast-WAM on matched A100 GPUs. Copper-Policy outperforms every compared method without embodied pretraining on RoboTwin and several embodied-pretrained VLAs on LIBERO-Plus (80.85%). On three challenging real-robot tasks, it performs comparably to $π_{0.5}$ and attains a higher average score. Together, these results show that Copper-Policy combines strong control performance with efficient training.