KeyGen:基于无监督关键点的对象中心表示用于类别级策略泛化
KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization
浏览论文内容
中文总结 AI 辅助
KeyGen通过无监督学习规范化3D关键点作为对象中心表示,结合扩散策略实现类别级操作泛化,在模拟和真实世界中均优于现有方法。
中文摘要 AI 辅助
机器人操作中的泛化要求策略能够跨各种形状、大小和姿态不同的未见对象实例执行任务。然而,传统的行为克隆(BC)方法往往过度拟合于实例特定的几何和外观,限制了向新对象的迁移。我们提出了KeyGen,一个从点云中学习规范化语义3D关键点并将其用作策略学习的结构化对象中心表示的框架。一个视觉运动扩散策略以这些关键点以及对象中心几何为条件,预测完整的操作轨迹,从而在对象实例间实现一致的几何对应。为了评估类别级泛化,我们构建了一个包含三个操作任务和规划驱动数据生成管线的逼真模拟基准,该管线在多样对象实例上生成专家轨迹。实验表明,KeyGen在姿态变化下对已见和未见对象均显著优于先前方法,随着每个对象的额外演示数量增加而有效扩展,对对象缩放保持鲁棒性,并在模拟和真实世界操作中均实现了强劲性能。
英文摘要
Generalization in robotic manipulation requires policies to perform tasks across diverse unseen object instances that vary in shape, size, and pose. However, conventional behavior cloning (BC) methods often overfit to instance-specific geometry and appearance, limiting transfer to novel objects. We introduce KeyGen, a framework that learns canonicalized semantic 3D keypoints from point clouds and uses them as structured object-centric representations for policy learning. A visuomotor diffusion policy conditions on these keypoints together with object-centric geometry to predict full manipulation trajectories, enabling consistent geometric correspondence across object instances. To evaluate category-level generalization, we construct a photorealistic simulation benchmark with three manipulation tasks and a planning-driven data generation pipeline that produces expert trajectories across diverse object instances. Experiments show that KeyGen significantly outperforms prior methods on both seen and unseen objects under pose variation, scales effectively with additional demonstrations per object, maintains robustness to object rescaling, and achieves strong performance in both simulation and real-world manipulation.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
- University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。