arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15631cs.RO

流匹配运动先验:用于模仿学习的在线最优传输奖励

Flow-Matched Motion Priors: Online Optimal-Transport Rewards for Imitation Learning

  • Tsinghua University(清华大学)
  • Chinese Academy of Sciences(中国科学院)

机构由 AI 辅助整理,请以论文原文为准。

Yilin Zou, Chenghua Liu, Chenglong Wu, Fanghua Jiang

AI总结:

针对模仿学习中对抗奖励在分布差异大时失效的问题,提出流匹配运动先验(FMP),利用最优传输耦合和流匹配神经势函数提供在线标量奖励,在Unitree G1上实现稳定行走并显著减少跌倒。

AI中文摘要:

学习运动先验需要一个奖励函数,引导策略从当前行为走向示范运动。对抗性运动先验(AMP)通过判别器提供这样的奖励。然而,当策略与专家支持分布相距较远时,对抗性目标可能变得信息量不足。朴素使用最优传输(OT)会将匹配的专家后继状态平均为重心目标。跨步态相位的平均会削弱目标的关节运动。我们提出流匹配运动先验(FMP),一种从连接当前轨迹历史与专家运动库的路径中学习到的在线标量奖励。熵正则化最优传输提供耦合。在每次策略更新前,我们沿轨迹到专家的路径,结合端点梯度监督和相对值校准,通过流匹配(FM)训练一个神经势函数。与AMP一样,actor仅接收物理观测,奖励保持为标量。受控的奖励模型实验表明,在拟合轨迹之外的泛化性能显著优于仅值函数或仅端点拟合。在Unitree G1上,匹配的5000万次转换实验比较了FMP与AMP、重心OT奖励以及嵌套消融,在示范和固定姿态初始化条件下。FMP在示范重置下产生0.727米/秒的稳定向前行走,在固定默认姿态下为0.338米/秒。在固定姿态条件下,其跌倒次数为129次,而仅端点控制为243次。与静态分数梯度教师相比,动态FM在插值分数0.25和0.50处降低了分数增量误差,同时离线拟合时间减少了29%。

英文摘要:

Learning a motion prior requires a reward that guides a policy from its current behavior toward demonstrated motion. Adversarial Motion Priors (AMP) provide such a reward with a discriminator. However, adversarial objectives can become uninformative when policy and expert supports are far apart. A naive use of optimal transport (OT) averages matched expert successors into a barycentric target. Averaging across gait phases can weaken the target's joint motion. We introduce Flow-Matched Motion Priors (FMP), an online scalar reward learned from paths connecting current rollout histories to an expert motion bank. Entropic OT supplies the coupling. Before each policy update, we train a neural potential with flow matching (FM) along the rollout-to-expert paths, endpoint-gradient supervision, and relative-value calibration. The actor receives only physical observations and the reward remains a scalar, as in AMP. Controlled reward-model experiments show substantially better generalization beyond the fitting rollout than value-only or endpoint-only fitting. On Unitree G1, matched 50-million-transition experiments compare FMP with AMP, a barycentric OT reward, and nested ablations under demonstration and fixed-pose initialization. FMP produces stable forward walking at 0.727 m/s from demonstration resets and 0.338 m/s from a fixed default pose. In the fixed-pose condition, it incurs 129 falls versus 243 for the endpoint-only control. Against a static score-gradient teacher, dynamic FM reduces score-increment error at interpolation fractions 0.25 and 0.50 while using 29% less offline fitting time.

↑