发表机构
Shanghai Jiao Tong University; Joy Future Academy, JD(上海交通大学; 京东探索研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
IronMan基于信息瓶颈原理,通过动力学感知瓶颈压缩视频特征,实现高效鲁棒的机器人操作,在LIBERO和RoboTwin上取得领先成功率并展现强OOD泛化能力。
AI 中文摘要
视频动作模型(VAMs)将视觉动力学建模与动作生成相结合,用于机器人操作。然而,视频表示并不天然适合动作生成,因为将动作策略暴露于过多的视觉细节会损害其泛化能力。因此,我们提出了IronMan(用于机器人操作的信息约束视频-动作学习),一个基于信息瓶颈原理构建的鲁棒视频-动作学习框架。该框架的核心原则是施加信息约束,以抑制无关的视觉信息,同时保留与动作相关的动力学线索。IronMan采用一种动力学感知瓶颈,将噪声纠缠的单步视频特征蒸馏为紧凑的世界表示。大量的仿真和真实世界实验证明了其在分布内(ID)性能上的强大表现和分布外(OOD)鲁棒性,同时保持了高效的推理。IronMan在LIBERO上达到了99.0%的成功率,在RoboTwin clean2clean上达到了79.4%的成功率,优于所有评估的基线方法。在OOD偏移下,IronMan在LIBERO-Plus上实现了79.1%的成功率,超过最强基线10.4个百分点。项目页面:此https URL。
英文摘要
Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation. However, video representations are not naturally suited to action generation, as exposing the action policy to excessive visual detail can impair its generalization ability. Therefore, we introduce IronMan (Information-constRained videO-actioN learning for robot MANipulation), a robust video-action learning framework built on the information bottleneck principle. The core principle of this framework is to impose information constraints that suppress irrelevant visual information while preserving action-relevant dynamics cues. IronMan employs a dynamics-aware bottleneck that distills noisy, entangled one-step video features into compact world representations. Extensive simulation and real-world experiments demonstrate strong in-distribution (ID) performance and out-of-distribution (OOD) robustness while maintaining efficient inference. IronMan achieves success rates of 99.0% on LIBERO and 79.4% on RoboTwin clean2clean, outperforming all the evaluated baselines. Under OOD shifts, IronMan achieves a success rate of 79.1% on LIBERO-Plus, exceeding the strongest baseline by 10.4 percentage points. Project page: https://youngsoul0731.github.io/ironman-project-page/
Comments21 pages, 12 figures, 5 tables