arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CriticHack:在机器人策略优化下评估视觉奖励

CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization

Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang

arXiv 2610.02527首次发表:更新:

发表机构

Baidu; University of California, Santa Barbara; Johns Hopkins University; Dalian University of Technology; Rutgers University(百度; 加州大学圣塔芭芭拉分校; 约翰霍普金斯大学; 大连理工大学; 罗格斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示视觉奖励模型优化会放大机器人策略的错误对象失败,提出倾斜模型解释该现象,并证明冻结结果验证器可纠正优化方向。

AI 中文摘要

学习得到的视觉奖励模型越来越多地被用于优化机器人策略,然而一个奖励模型可能将作用于错误对象的执行过程评为与完成任务同样高的分数。我们表明,优化这样的奖励会放大这些错误对象失败,同时奖励和任务成功率都在上升,因此从业者通常监控的信号看起来是健康的。我们在一个抽屉任务上,针对Robometer微调了扩散策略的每个去噪器参数。从没有先前奖励暴露的监督策略开始,五次训练运行在512个评估种子上将任务成功率提高了10.2个百分点,并将错误对象失败提高了10.9个百分点,而五次在模拟器的任务完成信号上训练的运行则在不放大错误对象失败的情况下提高了成功率(差异9.2个百分点,95%置信区间5.6至13.0)。这种放大现象在之前针对学习得到的奖励优化过的策略上,在该策略的原生扩散采样器下,在与初始策略匹配的距离处,以及在两个评论家和两个优化器的约束策略实验中反复出现。一个倾斜模型解释了其发生条件:在KL正则化优化下,当某个结果在初始策略下的期望奖励超过总体平均值时,该结果变得更加频繁。Robometer总体上很好地区分了成功和失败(AUROC 0.81),但将错误对象失败评分略高于成功(AUROC 0.37),因此优化提高了两者。同一模型预测了26个约束设置中的结果变化(Spearman 0.89),包括那些任务成功率下降的设置,并且Robometer自己发表的成功终止配方继承了这一错误。一个冻结的结果验证器将相同的优化重定向到所请求的任务。

英文摘要

Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task. We show that optimizing such a reward can amplify these wrong-object failures while reward and task success both rise, so the signals a practitioner would normally monitor look healthy. We fine-tune every denoiser parameter of a diffusion policy against Robometer on a drawer task. Starting from a supervised policy with no prior reward exposure, five training runs raise task success by 10.2 percentage points and wrong-object failures by 10.9 points on 512 evaluation seeds, whereas five runs trained on the simulator's task-completion signal raise success without amplifying wrong-object failures (difference 9.2 points, 95% CI 5.6 to 13.0). The amplification recurs from a policy previously optimized against learned rewards, under the policy's native diffusion sampler, at matched distance from the initial policy, and across constrained-policy experiments with two critics and two optimizers. A tilt model explains when it occurs: under KL-regularized optimization, an outcome becomes more frequent whenever its expected reward under the initial policy exceeds the population average. Robometer separates successes from failures well overall (AUROC .81) but scores wrong-object failures slightly above successes (AUROC .37), so optimization raises both. The same model predicts the outcome shifts across 26 constrained settings (Spearman .89), including those in which task success falls, and Robometer's own published success-termination recipe inherits the error. A frozen outcome verifier redirects the same optimization toward the requested task.

Comments60 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑