基于条件流匹配的高效多任务操作策略蒸馏
Distillation for Efficient Multitask Manipulation Policies via Conditional Flow Matching
浏览论文内容
中文总结 AI 辅助
本文提出通过转移单任务CFM专家的速度场,将知识蒸馏到共享多任务策略,结合原始CFM目标,在RLBench上以固定模型大小提升多任务性能。
中文摘要 AI 辅助
生成建模的进展最近被广泛用于机器人策略学习。特别是,使用专家演示训练的条件流匹配(CFM)已被证明在机器人操作基准上优于现有方法。虽然先前的工作主要关注单任务设置,我们从多任务的角度研究该问题,因为为每个任务训练独立模型在计算上是昂贵的。多任务策略学习有其自身的挑战,因为天真地在演示的拼接数据集上训练要么需要增加模型容量以适应增加的复杂性,要么导致性能下降。我们提出通过转移单任务CFM专家学到的速度场,将知识蒸馏到共享的多任务策略中。我们将这种蒸馏信号与原始CFM目标相结合,以保持对演示的保真度。在RLBench上的实验表明,我们的方法在保持固定模型大小的同时,提高了多任务策略性能,优于天真训练。
英文摘要
Advances in generative modeling have recently been extensively employed in robotics for policy learning. In particular, Conditional Flow Matching (CFM) trained with expert demonstrations has been shown to outperform existing methods on robot manipulation benchmarks. While prior work has mainly focused on single-task settings, we study the problem from a multi-task perspective, as training independent models for each task is computationally expensive. Multi-Task policy learning comes with its own set of challenges, as naively training on a concatenated dataset of demonstrations would either require increased model capacity to accommodate the added complexity or result in drops in performance. We propose to distill knowledge from single-task CFM experts into a shared multi-task policy by transferring their learned velocity fields. We combine this distillation signal with the original CFM objective to retain fidelity to the demonstrations. Experiments on RLBench show that our approach improves multi-task policy performance over naive training while maintaining a fixed model size.
发表机构
- University of Freiburg(弗莱堡大学)
机构由 AI 辅助整理,请以论文原文为准。