arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06852cs.RO

ContextFlow: 机器人操作中的上下文流匹配

ContextFlow: In-Context Flow Matching for Robot Manipulation

Jian Ding, Xianjie Dai, Roei Herzig, Nussair Hroub, Jinjie Mai, Dengxin Dai, Bernard Ghanem, Mohamed Elhoseiny

首次发表
浏览论文内容

中文总结 AI 辅助

ContextFlow提出条件流匹配模型,利用上下文演示学习连续动作分布,在LIBERO上超越ICRT 35个百分点,并无需微调匹配π0性能,在真实机器人上泛化到未见配置。

中文摘要 AI 辅助

尽管在视觉和语言领域非常有效,但将上下文学习应用于机器人技术仍然具有挑战性。现有的自回归上下文模仿方法将连续动作离散化,并通过下一个词元预测加剧了早期预测误差的累积,限制了它们在未见任务配置上的泛化能力。同时,流匹配策略已被探索用于连续机器人控制,并有助于缓解复合误差;然而,在流匹配框架内的上下文模仿学习仍未得到充分探索。为了解决这些局限性,我们引入了ContextFlow,一种条件流匹配模型,用于学习连续动作分布以进行上下文模仿学习。ContextFlow将基于流的动作预测条件于演示和观察,从而能够从噪声动作分布中稳健生成。为了更好地编码多模态上下文演示,我们采用了感知器风格的多模态上下文压缩器,将视觉、本体感觉和动作序列蒸馏为紧凑的、与任务相关的潜在表示。在LIBERO上,ContextFlow在未见任务配置上的平均成功率比ICRT高出35个百分点,同时在未见任务上无需任何微调即可匹配任务特定微调的VLA模型π0的性能。在真实机器人上,它泛化到单臂和双臂任务的未见配置,在新的笔帽开启配置上实现了40%的成功率。项目页面:此https URL。

英文摘要

Although highly effective in vision and language domains, applying in-context learning to robotics remains challenging. Existing autoregressive in-context imitation methods discretize continuous actions and exacerbate the accumulation of early prediction errors through next-token prediction, limiting their generalization on unseen task configurations. Meanwhile, flow-matching policies have been explored for continuous robot control and can help mitigate compounding errors; however, in-context imitation learning within a flow-matching framework remains underexplored. To address these limitations, we introduce ContextFlow, a conditional flow-matching model that learns continuous action distributions for in-context imitation learning. ContextFlow conditions flow-based action prediction on demonstrations and observations, enabling robust generation from noisy action distributions. To better encode multimodal in-context demonstrations, we adapt perceiver-style multimodal context compressors that distill visual, proprioceptive, and action sequences into compact, task-relevant latent representations. On LIBERO, ContextFlow outperforms ICRT by 35 percentage points in average success rate on unseen task configurations, while matching the performance of the task-specific fine-tuned VLA model $π_0$ without any fine-tuning on unseen tasks. On real robots, it generalizes to unseen configurations of both single-arm and bimanual tasks, achieving 40% success on a new pen-uncapping configuration. Project Page: https://dingjiansw101.github.io/contextflow-page/.

发表机构

  • KAUST(阿卜杜拉国王科技大学)
  • University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑