发表机构
Peking University; AI Robotics(北京大学; AI Robotics)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RFPO通过奖励感知在线Reflow和冻结高斯PPO监督,使流策略在少步执行下保持全步性能,在多种机器人上实现一步推理且性能损失小于2.4%,推理速度提升54.9倍。
AI 中文摘要
基于流的策略为连续机器人控制提供了富有表现力的框架,但其迭代式常微分方程(ODE)积分会带来可观的推理成本。天真地减少积分预算可能会严重降低控制性能,因为在全步执行下优化的策略并未被明确约束为在粗粒度数值积分下仍然可靠。我们将这种不匹配称为少步离散化差距。为解决此问题,我们引入了RFPO,一个用于可靠少步执行的流策略优化框架。奖励感知的在线Reflow在策略学习过程中矫正学生诱导的传输路径,使所得策略对粗粒度积分更加鲁棒。一个冻结的高斯PPO控制器在全积分和中间积分预算下提供补充的动作空间监督,而部署的策略仍然是仅用一个欧拉步执行单个流学生。在Unitree Go2、Boston Dynamics Spot、Unitree H1和Unitree G1上,RFPO在一步执行下始终保留全步控制性能,一步回报在零初始化和随机初始化下均保持在相应64步值的2.4%以内。在Unitree Go2上,一步执行保留了64步奖励的98.5%,同时将板载平均推理延迟从4.39毫秒降低到0.08毫秒,实现了54.9倍的加速。真实机器人实验进一步验证了稳定的一步行走。代码:此https URL。网站:此https URL。
英文摘要
Flow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost. Naively reducing the integration budget can severely degrade control, since policies optimized under full-step execution are not explicitly constrained to remain reliable under coarse numerical integration. We refer to this mismatch as the few-step discretization gap. To address this problem, we introduce RFPO, a flow-policy optimization framework for reliable few-step execution. Reward-aware online Reflow rectifies student-induced transport paths during on-policy learning, making the resulting policy more robust to coarse integration. A frozen Gaussian PPO controller supplies complementary action-space supervision at full and intermediate integration budgets, while the deployed policy remains a single flow student executed with one Euler step. Across Unitree Go2, Boston Dynamics Spot, Unitree H1, and Unitree G1, RFPO consistently preserves full-step control performance under one-step execution, with one-step returns remaining within 2.4% of their corresponding 64-step values across both zero and random initialization. On Unitree Go2, one-step execution retains 98.5% of the 64-step reward while reducing onboard mean inference latency from 4.39 ms to 0.08 ms, yielding a 54.9x speedup. Real-robot experiments further validate stable one-step locomotion. Code: https://github.com/AIGeeksGroup/RFPO. Website: https://aigeeksgroup.github.io/RFPO.