发表机构
Toyota Research Institute; Woven by Toyota; Cornell University(丰田研究院; 丰田编织公司; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究微调大型行为模型时先验选择的影响,发现非高斯先验相比标准高斯先验在多数情况下无显著优势,编码器训练是主导因素。
AI 中文摘要
现代机器人模仿学习日益依赖于基于扩散或流匹配模型的生成式策略,这些模型通过变换来自先验分布的样本来生成动作。一个关键问题是先验的选择是否重要。已有研究表明,在从头训练时,用更接近目标的非高斯先验替代标准高斯分布可显著提升性能。一个自然的后续问题是,这些收益能否迁移到微调预训练的大型行为模型(LBM),如LBM 1.0、π₀.₅和GR00T N1.5,在这些场景中人们可能预期收益更大。令人惊讶的是,我们发现情况并非如此,除非在极低的微调数据比例下。在跨越上述三种LBM、涉及两个仿真平台中40多项任务、超过10万次仿真 rollout,以及五项双臂操作任务上的1250次硬件 rollout中,那些被证明更接近目标的非高斯先验在微调性能上与标准高斯先验在统计上无显著差异或更差。诊断分析揭示了原因:微调后的模仿学习策略在不同先验下收敛到相似的动作预测,尽管其微调后的编码器嵌入与预训练嵌入及彼此之间差异显著。一项学习率消融实验进一步证实,编码器训练是微调性能的主导因素,其影响远超先验选择。我们最后为未来关于学习先验在微调中何时以及为何仍可能起作用的研究提出了具体方向。项目页面:https://cxu-tri.github.io/non_gaussian_FT/
英文摘要
Modern robot imitation learning increasingly relies on generative policies based on diffusion or flow-matching models, which generate actions by transforming samples from a prior distribution. A key question is whether the choice of prior matters. Replacing the standard Gaussian with a closer-to-target, non-Gaussian prior has been shown to substantially improve performance when training from scratch. A natural next step is to ask whether these gains transfer to fine-tuning pretrained Large Behavior Models (LBMs) such as LBM 1.0, $π_{0.5}$, and GR00T~N1.5, where one might expect even larger gains. Surprisingly, we find that this is not the case, except possibly at very low fine-tuning data fractions. Across over 100K simulation rollouts spanning all three aforementioned LBMs on 40+ tasks in two simulation platforms, and 1250 hardware rollouts on five bimanual manipulation tasks, non-Gaussian priors that are demonstrably closer to the target yield statistically indistinguishable or worse fine-tuning performance than a standard Gaussian prior. Diagnostic analyses suggest why: fine-tuned imitation learning policies converge to similar action predictions across priors, despite their fine-tuned encoder embeddings diverging substantially from the pretrained embeddings and each other. A learning-rate ablation further confirms that encoder training is the dominant factor in fine-tuning performance, substantially outweighing the effect of prior choice. We conclude with concrete directions for future research on when and why learned priors might still matter in fine-tuning. Project page: https://cxu-tri.github.io/non_gaussian_FT/