发表机构
Korea Advanced Institute of Science & Technology (KAIST)(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EquivDP3通过将SIM(3)等变编码器集成到分层扩散策略中,在少量演示下显著提升人形机器人移动操作的成功率,且几乎不增加计算开销。
AI 中文摘要
用于人形机器人移动操作的视觉运动策略必须仅从少量演示中泛化到不同的物体姿态和光照条件。3D扩散策略(DP3)以点云特征为条件生成扩散动作,但其基于PointNet风格的编码器缺乏对操作任务所遵循的旋转、平移和缩放(SIM(3))的内建等变性。EquiBot通过使用SIM(3)等变向量神经元网络(VNN)编码器弥补了轮式操作机器人的这一不足。我们将此扩展到更为复杂的实体——具有43个关节的Unitree G1人形机器人,并提出EquivDP3:一种两阶段策略,其中高层扩散规划器使用SIM(3)等变VNN编码器以6Hz频率生成全身命令块,由冻结的预训练强化学习运动策略和用于手臂的差分逆运动学模块以50Hz频率执行,并通过行为克隆进行端到端训练。在两个模拟的IsaacLab基准测试和四个非等变基线(5-100个演示,分布内和分布外)中,EquivDP3的优势集中在低数据区域:在5-10个演示时,其成功率可达67.1%,而基线为38.2%-52.4%;到50-100个演示时,所有编码器均收敛(74.3%-85.2%),排序不再具有意义。仅使用本体感觉的对照实验证实了这一差距确实源于感知:移除点云后,在5个演示时成功率从EquivDP3的60%降至31%,在10个演示时从99%降至78%,但在50-100个演示时差距消失,表明高数据平台反映了基准上限,而非五种编码器学习了相同的不变性。该编码器相比其替换的PointNet编码器,每个动作块仅增加0.8毫秒的额外延迟。将几何对称性融入分层扩散策略的感知主干,是在演示稀缺时提高人形机器人移动操作数据效率的一种实用且几乎零成本的方法。
英文摘要
Visuomotor policies for humanoid loco-manipulation must generalize across object poses and lighting from only a handful of demonstrations. 3D Diffusion Policy (DP3) conditions a diffusion-based action generator on point-cloud features, but its PointNet-style encoder has no built-in equivariance to the rotations, translations, and scalings (SIM(3)) that manipulation tasks respect. EquiBot closed this gap for wheeled manipulators with a SIM(3)-equivariant Vector Neuron Network (VNN) encoder. We extend this to a substantially more complex embodiment, the 43-joint Unitree G1 humanoid, and propose EquivDP3: a two-stage policy where a high-level diffusion planner with a SIM(3)-equivariant VNN encoder emits 6 Hz whole-body command chunks, executed at 50 Hz by a frozen, pre-trained RL locomotion policy and a differential inverse-kinematics module for the arms, trained end-to-end by behavior cloning. Across two simulated IsaacLab benchmarks and four non-equivariant baselines (5-100 demonstrations, in- and out-of-distribution), EquivDP3's advantage concentrates in the low-data regime: at 5-10 demonstrations it reaches 67.1% success versus 38.2-52.4% for the baselines, while by 50-100 all encoders converge (74.3-85.2%) and the ordering is no longer meaningful. A proprioception-only control confirms this gap is genuinely perceptual: with the point cloud removed, success drops to 31% vs. 60% (EquivDP3) at 5 demonstrations and 78% vs. 99% at 10, but vanishes by 50-100, showing the high-data plateau reflects a benchmark ceiling, not five encoders learning the same invariance. The encoder costs only 0.8 ms of extra latency per action chunk over the PointNet encoder it replaces. Baking geometric symmetry into a hierarchical diffusion policy's perception backbone is a practical, nearly free way to improve data efficiency for humanoid loco-manipulation when demonstrations are scarce.
Comments12 pages, 5 figures, 5 tables. A shorter version has been submitted to ICASSP 2027