arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40165cs.ROcs.AIcs.LG

PrefPI:偏好引导转向分布外行为

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

首次发表
浏览论文内容

中文总结 AI 辅助

PrefPI利用相对偏好通过偏好条件生成建模和无分类器引导迭代优化预训练机器人策略,实现分布外行为转向,在真实硬件上以150条偏好轨迹将运输高度提升至19.8厘米。

中文摘要 AI 辅助

我们提出PrefPI(偏好引导策略迭代),一种迭代框架,仅利用自生成轨迹上的相对偏好来引导预训练的生成式机器人策略。与先前主要锐化策略已表示模式的偏好学习方法不同,我们研究超越初始有效支持的转向,其中期望行为在初始策略下很少或从未被观察到。我们的关键思想是将偏好学习表述为偏好条件生成建模:偏好轨迹定义了一个条件分布,其与更广泛行为先验的密度比提供了一个隐式偏好信号,并通过无分类器引导(CFG)放大。重复这种偏好条件建模和引导步骤产生了一种偏好引导策略迭代形式,将增量改进转向先前不可达的行为。在扩散策略和PI0.5流匹配VLA的仿真和真实世界中,PrefPI以有限的反馈产生了显著的行为转变。特别是,PrefPI在真实硬件上仅用150条偏好标记轨迹就将物体运输高度从10.7厘米提高到19.8厘米。

英文摘要

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering beyond the initial effective support, where desired behaviors are rarely or never observed under the initial policy. Our key idea is to formulate preference learning as preference-conditioned generative modeling: preferred trajectories define a conditional distribution, whose density ratio with the broader behavior prior provides an implicit preference signal amplified by classifier-free guidance (CFG). Repeating this preference-conditioned modeling and guidance step yields a form of preference-guided policy iteration, turning incremental improvements toward previously inaccessible behaviors. Across diffusion policies and the PI0.5 flow- matching VLA in simulation and the real world, PrefPI produces substantial behavioral shifts with limited feedback. In particular, PrefPI increases object transport height from 10.7 cm to 19.8 cm on real hardware with only 150 preference-labeled trajectories.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑