发表机构
Gaoling School of Artificial Intelligence, Renmin University of China; Beijing Innovation Center of Humanoid Robotics; Beijing Forestry University; Peking University; Microsoft Research(中国人民大学高瓴人工智能学院; 北京人形机器人创新中心; 北京林业大学; 北京大学; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究离线到在线强化学习中价值函数可靠性对策略优化的影响,提出Robo-ValueRL框架,通过特定指标评估价值估计可靠性并用于策略预训练和在线改进,实验显示其能有效提升下游策略性能,突出价值引导数据利用对策略改进的重要性。
AI 中文摘要
离线到在线强化学习在可推广的机器人操纵方面很有前景,但其全栈复杂性影响了再现和诊断。价值估计在为策略改进确定异构数据优先级方面起着核心作用,但价值函数可靠性如何影响离线到在线强化学习中的策略优化这一核心问题仍未得到充分探索。为此提出Robo-ValueRL统一框架,通过全局进展和局部偏好指标评估价值估计器可靠性,将其用于策略预训练和在线改进。实验表明下游性能与价值可靠性密切相关,可靠价值函数能提供更好的动作质量估计,使系统在芯片插入和块拆卸任务中取得高成功率。
英文摘要
Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems, value estimation plays a central role in prioritizing heterogeneous data for policy improvement. Despite its importance, the central question remains underexplored: how value-function reliability shapes policy optimization in offline-to-online reinforcement learning. To answer this question, we propose Robo-ValueRL, a unified framework that enables reliable value estimation and systematically traces its downstream effects on policy pretraining and online improvement. Concretely, Robo-ValueRL learns a history-conditioned value estimator and evaluates its reliability through global-progress and local-preference metrics. These resulting value estimates are propagated into quality-conditioned consistency-policy pretraining and a residual adaptation module on online rollouts, providing a unified testbed for analyzing how value reliability shapes downstream policy performance. Across 240 hours of offline demonstrations and over 3,000 online rollout trajectories, our extensive experiments show that downstream performance is strongly associated with value reliability. Reliable value functions provide better action-quality estimates, allowing value-guided offline RL to scale more effectively than quality-agnostic behavior cloning, and stabilize online improvement by prioritizing high-quality rollout data. Integrating reliable value guidance through offline pretraining with online improvement, our system achieves 86% success on millimeter-level precise chip insertion and 84% on generalizable block disassembly. We hope these findings highlight the importance of value-guided data utilization for effective policy improvement from heterogeneous robotic experience.
CommentsPlease refer to our website: https://gewu-lab.github.io/Robo-ValueRL/