arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36872cs.RO

PreferenceFlow:基于人类干预的流匹配机器人策略测试时引导

PreferenceFlow: Test-Time Guidance of Flow-Matching Robot Policies from Human Interventions

Yiqi Tang, Diyuan Shi, Runze Li, Donglin Wang

首次发表
浏览论文内容

中文总结 AI 辅助

PreferenceFlow利用人类干预片段训练偏好模型,在测试时引导流匹配机器人策略,无需奖励信号,在精密插入任务中成功率从69%提升至90.5%。

中文摘要 AI 辅助

流匹配策略能够表示复杂的机器人行为,但在部署时面临分布偏移,容易产生局部错误。许多用于策略改进的强化学习方法需要奖励信号,而在真实世界的操作任务中,这些信号难以指定或获取。我们提出了PreferenceFlow,一个在测试时改进预训练流策略的框架,无需环境奖励或对基础策略进行更新。将人类干预片段与从相同初始条件状态生成的机器人片段配对,用于训练偏好模型。在推理过程中,我们采用QGF采样更新,将其值梯度替换为在估计的干净动作处评估的偏好梯度。梯度上限损失对干预对上的过大梯度进行惩罚,而零梯度损失则抑制在专家演示动作附近的引导。在Franka机器人上进行的四项真实世界精密插入任务中,PreferenceFlow实现了90.5%的平均成功率,而冻结策略的成功率为69%。消融实验支持了梯度正则化和正确排序的偏好标签在所评估设置中的作用。这些结果证明了人类干预作为局部偏好监督在引导冻结生成式机器人策略方面的有效性。

英文摘要

Flow-matching policies can represent complex robot behaviors but remain susceptible to local errors under distribution shift at deployment. Many reinforcement learning approaches to policy improvement require reward signals that are difficult to specify or obtain in real-world manipulation. We present PreferenceFlow, a framework for improving a pretrained flow policy at test time without environment rewards or updates to the base policy. Human intervention chunks are paired with robot chunks generated from the same initial conditioning state to train a preference model. During infer- ence, we adopt the QGF sampling update, replacing its value gradient with the preference gradient evaluated at an estimated clean action. A gradient-cap loss penalizes excessive gradients on intervention pairs, while a zero-gradient loss discourages guidance near actions from expert demonstrations. On four real-world precision insertion tasks with a Franka robot, Pref- erenceFlow achieves a mean success rate of 90.5%, compared with 69% for the frozen policy. Ablations support the roles of gradient regularization and correctly ordered preference labels in the evaluated settings. These results demonstrate the utility of human interventions as local preference supervision for guiding frozen generative robot policies.

发表机构

  • Westlake University(西湖大学)

机构由 AI 辅助整理,请以论文原文为准。

↑