arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00864cs.ROcs.AIcs.CLcs.CV

运动学均值流:面向机器人基础模型的一步动作生成策略

Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Kinematic MeanFlow(K-MF),一种面向机器人基础模型的一步动作生成策略,通过解耦时间导数项解决MeanFlow性能崩溃问题,在多数任务上优于多步流匹配,并显著降低推理延迟。

中文摘要 AI 辅助

在本文中,我们研究如何在机器人基础模型(RFMs)中实现一步动作生成,旨在克服多步流匹配的高推理延迟。MeanFlow为此目标提供了一个有前景的框架,然而其直接应用会导致性能崩溃。我们发现这源于RFM速度场中表现出的两种独特动态:(1)“局部加速度”在早期保持稳定,但在去噪过程结束时急剧飙升;(2)其幅度在样本间的分布随去噪进行而扩大。为解决这些问题,我们提出了Kinematic MeanFlow(K-MF),一种专为RFMs设计的新型一步动作策略。具体而言,基于运动学恒等式,K-MF将MeanFlow公式中的时间导数项分解为两个由中间点分隔的子区间项。这种解耦公式使得这两项能够分别捕获早期和晚期去噪动态,同时减轻整个过程中的误差放大。因此,我们的K-MF使RFMs能够在从头训练和微调范式中实现跨多样任务的一步动作生成,并在大多数设置中优于多步流匹配。在推理效率方面,K-MF在L40和Jetson Orin上以急切和编译模式将GR00T-N1.6的动作头延迟降低了67.5%~74.4%,实现了端到端延迟降低30.3%~54.9%。代码将在此https URL提供。

英文摘要

In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems from two distinctive dynamics exhibited in the RFM velocity field: (1) the ``local acceleration" exhibits stability early on, but surges sharply towards the end of the denoising process, and (2) the spread of its magnitudes across samples widens as denoising progresses. To address these issues, we introduce Kinematic MeanFlow (K-MF), a novel one-step action policy tailored for RFMs. Specifically, grounded in a kinematic identity, K-MF decouples the time derivative term in the MeanFlow formulation into two sub-interval terms separated by an intermediate point. This decoupled formulation enables the two terms to capture early-stage and late-stage denoising dynamics, respectively, while mitigating the error amplification across the process. As a result, our K-MF empowers RFMs to achieve one-step action generation in both training from scratch and fine-tuning paradigms across diverse tasks, while outperforming multi-step flow matching in most settings. In terms of inference efficiency, K-MF reduces action-head latency of GR00T-N1.6 by 67.5%~74.4% across L40 and Jetson Orin in eager and compiled modes, yielding end-to-end latency reductions of 30.3%~54.9%. Code will be available at https://github.com/IntelChina-AI/K-MF.

发表机构

  • Intel Labs China(英特尔中国研究院)
  • Midea AI Research(美的AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑