arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

QF3:基于滤波Q梯度的快速流强化学习

QF3: Fast Flow RL with Filtered Q-Gradients

Chung Min Kim, Brent Yi, David McAllister, Hongsuk Choi, Himanshu Gaurav Singh, Jinkun Cao, Ken Goldberg, Pieter Abbeel, Carmelo Sferrazza, Angjoo Kanazawa

arXiv 2610.08789首次发表:更新:

发表机构

UC Berkeley; Amazon FAR(加州大学伯克利分校; 亚马逊FAR)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

QF3是一种离策略流强化学习算法,通过滤波Q梯度训练流策略,实现从零训练人形运动策略并零样本迁移到硬件,速度比FPO++快10倍,且能微调预训练操作策略。

AI 中文摘要

流策略已成为从示范中学习机器人行为的标准策略类别,但强化学习对于改进预训练的流策略或通过交互从零开始学习它们仍然至关重要。我们引入了QF3(基于滤波Q梯度的快速流强化学习),一种在线离策略强化学习算法,该算法使用流匹配加上评论家的动作梯度来训练流策略,该梯度通过流输出的一步预测进行反向传播。为了将更新保持在评论家和该预测可靠的区域,QF3仅将评论家梯度应用于靠近重放动作的动作维度。据我们所知,QF3是第一种从零开始训练人形运动策略并将其零样本迁移到硬件的离策略流强化学习方法。与高通量离策略训练方案相结合,它在人形运动和运动跟踪策略上的训练速度比FPO++(一种最近的在线策略流强化学习方法)快10倍。我们进一步将QF3应用于在ABC-Sim和Robomimic任务上微调预训练的基于流的操作策略。这些结果表明,QF3既能从零开始学习机器人策略,也能改进从示范中获得的策略。网站:此https URL

英文摘要

Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic's action gradient, backpropagated through a one-step prediction of the flow's output. To keep updates where the critic and this prediction are reliable, QF3 applies the critic gradient only to action dimensions that stay near the replay action. To our knowledge, QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to hardware. Paired with a high-throughput off-policy training recipe, it trains humanoid locomotion and motion-tracking policies with a 10x wall-clock speedup over FPO++, a recent on-policy flow RL method. We further apply QF3 to fine-tune pretrained flow-based manipulation policies on both ABC-Sim and Robomimic tasks. These results suggest that QF3 can both learn robot policies from scratch and refine those acquired from demonstrations. Website: https://qf3-rl.github.io/

CommentsProject page: https://qf3-rl.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑