OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
OGPO:生成控制策略的样本高效全微调
机构 * University of California, Berkeley(加州大学伯克利分校) ; UC Berkeley(加州大学伯克利分校)
AI总结 提出OGPO算法,通过离策略评论网络和修改的PPO目标,实现生成控制策略的样本高效微调,在多种操作任务上达到最优性能,并能在无专家数据下微调不良初始化的行为克隆策略。
Comments Website: https://simchowitzlabpublic.github.io/ogpo-site/ Code: https://github.com/simchowitzlabpublic/OGPO_public