发表机构
University of Pennsylvania; Tsinghua University; UC San Diego(宾夕法尼亚大学; 清华大学; 加利福尼亚大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对机器人学习中回归策略与生成策略的性能差距,提出异方差学生t动作回归策略HT-Policies,其训练推理更快,在仿真与真实机器人任务中性能可与生成策略基线媲美。
AI 中文摘要
从演示中学习已能实现令人印象深刻的机器人行为。策略学习的常用选择是使用扩散或流匹配(Flow-Policies),其通常优于用均方误差(MSE)训练的直接动作回归策略(MSE-Policies)。这一差距通常归因于演示的多模态特性。我们从统计建模的角度重新审视这一差距:动作预测残差如何影响策略优化。对真实世界机器人演示数据的分析显示,残差尺度存在显著的状态依赖变化,且具有比高斯分布更重的尾部。尽管MSE-Policies和Flow-Policies均呈现重尾动作残差,但它们的训练梯度表现不同:MSE会将更多梯度幅度分配给动作残差大的观测值,这会损害优化效果。基于这些发现,我们引入异方差学生t动作回归策略(HT-Policies),该策略可学习输入依赖的残差尺度并降低重尾的影响。HT-Policies通过单次前馈传递预测动作块,且可复用预训练的基于流匹配的策略网络作为骨干。在四个仿真基准测试和真实机器人评估中,HT-Policies在从头训练和从预训练的视觉-语言-动作及世界-动作模型训练时,均达到了与生成策略基线相当的成功率,同时训练和推理速度更快。这些发现阐明了生成目标在从演示学习的机器人学习中的实际优势,并为一系列架构和任务提供了高效的直接回归替代方案。项目页面:this https URL
英文摘要
Learning from demonstration has enabled impressive robot behaviors. A common choice for policy learning is to use diffusion or flow matching (Flow-Policies), which often outperforms direct action regression trained with mean squared error (MSE-Policies). This gap is commonly attributed to multimodal demonstrations. We revisit this gap from the perspective of statistical modeling: how action-prediction residuals shape policy optimization. Our analysis of real-world robot demonstration data reveals substantial state-dependent variation in residual scales and heavier-than-Gaussian tails. While both MSE-Policies and Flow-Policies exhibit heavy-tailed action residuals, their training gradients behave differently: MSE allocates more gradient magnitude to observations with large action residuals, which hurts optimization. Motivated by these findings, we introduce heteroscedastic Student-t action regression (HT-Policies), which learns input-dependent residual scales and reduces the influence of heavy tails. HT-Policies predict action chunks with a single feed-forward pass and can reuse pretrained flow-matching-based policy networks as the backbone. Across four simulation benchmarks and real-robot evaluations, HT-Policies achieves success rates competitive with generative policy baselines, both when trained from scratch and from pretrained vision-language-action and world-action models, despite being faster in training and inference. Together, these findings shed light on the practical advantages of generative objectives in robot learning from demonstrations and offer an efficient direct-regression alternative for a range of architectures and tasks. Project page: https://the-labone.github.io/regression-policy-project/