arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双线性流策略:目标条件视觉运动模仿的分布外推

Bilinear Flow Policy: Distributional Extrapolation for Goal-Conditioned Visuomotor Imitation

Wonsuhk Jung, Sundhar Vinodh Sangeetha, Chen Xu, Abhishek Gupta, Masha Itkina, Shreyas Kousik, Haruki Nishimura

arXiv 2610.05765首次发表:更新:

发表机构

Georgia Institute of Technology; Toyota Research Institute; University of Washington(佐治亚理工学院; 丰田研究所; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出双线性流策略(BFP),通过转导检索和双线性条件流实现目标条件模仿学习的分布外推,在模拟和真实任务中显著提升未见目标的成功率。

AI 中文摘要

目标条件模仿学习(GCIL)结合流匹配是一个有前景的框架,能够表示多模态行为并适应多样化的用户指定目标,但当目标位于演示支持范围之外时,该框架常常失败。为了在不破坏多模态性的情况下外推到此类未见目标——我们将这一问题称为分布外推——我们引入了双线性流策略(BFP),这是一种生成式视觉运动策略,将转导检索与双线性条件流相结合。给定一个未见观测-目标对,BFP检索一个“锚点”训练示例,并将该未见对转导性地重构为这个熟悉的锚点加上一个残差项。为了使这种分解能够指导动作预测,残差必须紧凑地编码当前观测-目标对与锚点之间的差异,并且锚点的选择必须使得该差异能够预测相应的动作分布。BFP通过预训练视觉特征和一种新颖的学习式锚点选择算法实现这一点。新颖的双线性流随后建模锚点和残差如何共同决定多模态动作分布。我们证明,在适当假设下,对于双线性流,未见目标处的动作分布误差受分布内流匹配误差乘以问题相关因子的限制。在模拟中的五个操作任务上,BFP的分布外成功率是GCIL策略的2.63倍,是最强外推定向基线的1.36倍。在两个真实世界任务上,BFP相对于GCIL提升了32%。最后,我们的理论为预测哪些训练好的策略能够良好外推以及外推到哪些未见目标提供了实用的部署前诊断方法。

英文摘要

Goal-conditioned imitation learning (GCIL) with flow matching is a promising framework that can represent multimodal behaviors while adapting to diverse, user-specified goals, yet often fails when goals lie outside the demonstration support. To extrapolate to such unseen goals without collapsing multimodality - a problem we call distributional extrapolation - we introduce Bilinear Flow Policy (BFP), a generative visuomotor policy that combines transductive retrieval with a bilinear conditional flow. Given an unseen observation-goal pair, BFP retrieves an "anchor" training example and transductively reformulates the unseen pair as this familiar anchor plus a residual term. For this decomposition to guide action prediction, the residual must compactly encode how the current observation-goal pair differs from the anchor, and the anchor must be chosen so that this difference is predictive of the corresponding action distribution. BFP achieves this with pretrained visual features and a novel learned anchor-selection algorithm. The novel bilinear flow then models how the anchor and the residual jointly determine the multimodal action distribution. We prove that, for bilinear flow under suitable assumptions, action distribution error at unseen goals is bounded by the in-distribution flow-matching error up to problem-dependent factors. Across five manipulation tasks in simulation, BFP achieves 2.63x the out-of distribution success rate of a GCIL policy and 1.36x that of the strongest extrapolation-targeted baseline. On two real-world tasks, BFP improves over GCIL by 32%. Finally, our theory yields practical, pre deployment diagnostics for predicting which trained policies will extrapolate well and to which unseen goal.

Comments40 pages, 18 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑