发表机构
KAIST; RLWRLD(韩国科学技术院; RLWRLD)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对流策略微调中伴随匹配计算成本高的问题,提出标量伴随匹配的Q学习(SQAM),利用速度雅可比对角化特性简化计算,并在多个难任务及真实机器人上显著提升性能。
AI 中文摘要
流策略能够捕获丰富多样的动作分布,利用离策略强化学习对其进行微调以超越示范,已引起越来越多的关注。然而,针对学习到的价值函数微调流策略并非易事,因为该策略在多个流步骤中生成其动作。伴随匹配提供了一种原则性方法,通过将最终动作的价值信息传播回每个流步骤来更新流模型本身,但它在每一步都需要通过策略进行向量-雅可比乘积,其成本随流步骤数和策略规模的增长而增加。我们观察到,预训练流策略的批量平均速度雅可比矩阵集中在其对角线上。受此发现启发,我们推导出一个闭式标量伴随,它将最终动作处的价值梯度按流时间缩放,从而消除了逐步骤的向量-雅可比乘积。我们进一步发现,在标量伴随下,控制评论家在策略生成动作处的价值尤为重要。基于这些发现,我们提出了标量伴随匹配的Q学习(SQAM),它将标量伴随与这些动作处的价值惩罚相结合。SQAM的收益集中在OGBench中最难的四个领域,其成功率在各自领域超过最强基线的18至35个百分点。为了测试SQAM是否适用于大型预训练策略,我们还在真实双臂机器人上微调了一个视觉-语言-动作策略。SQAM在所有三个任务上都优于监督微调。
英文摘要
Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products. We further find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.