AI 中文总结
研究如何使机器人策略与人类偏好对齐,提出基于偏好的奖励聚类(PREC)框架,通过聚合用户轨迹学习共享轨迹编码器,将用户聚类并为每个聚类学习奖励模型,实验表明该方法能更准确聚类用户,在多指标上优于现有方法。
AI 中文摘要
使机器人策略与人类偏好对齐对于部署到不同终端用户至关重要。在逐用户对齐方法中,偏好反馈稀疏,学习不稳定且易受人类偏好噪声影响,大量个性化策略在部署前难以验证。单一共享策略方法虽避免成本,但无法捕捉异质偏好且常忽略少数群体偏好。为应对这些挑战,我们引入基于偏好的奖励聚类(PREC)框架,从不同用户提供的二元偏好标签中学习一组紧凑的策略。PREC先将标签搁置一旁,聚合用户轨迹以学习群体级共享轨迹编码器,减轻每个用户覆盖范围有限的问题并避免表示学习中的标签噪声。利用此表示,PREC将用户联合分配到偏好一致的聚类中,并为每个聚类学习代表性奖励模型,从中为每个聚类优化策略。聚类相似用户弥补了每个用户可用标签数量有限的问题并减轻标签噪声影响。同时,维持可管理数量的奖励模型减轻了部署时的验证负担。在不同模拟运动环境中的实验表明,PREC比基线方法更准确地将标记不同轨迹子集的用户分组到偏好一致的聚类中。在稀疏和噪声反馈下,用PREC训练的策略在所有三个社会福利指标上优于现有的单一共享策略用户对齐方法,甚至超过逐用户对齐方法。
英文摘要
Aligning robot policies with human preferences is essential for deployment to diverse end users. In per-user alignment approach, preference feedback is often sparse, so learning becomes unstable and vulnerable to human preference noise, and a growing number of individualized policies makes validation difficult before deployment. A single shared policy approach to user alignment avoids this cost but fails to capture heterogeneous preferences and often neglects minority preferences. To address these challenges, we introduce Preference-based REward Clustering (PREC), a novel framework that learns a compact set of policies from binary preference labels provided by diverse users. From a dataset of user trajectories and their preference labels, PREC first sets the labels aside and aggregates trajectories across users to learn a population-level shared trajectory encoder, alleviating limited per-user coverage and avoiding label noise during representation learning. Using this representation, PREC jointly assigns users to preference-coherent clusters and learns a representative reward model per cluster using preference labels, from which a policy is optimized for each cluster. Clustering similar users compensates for the limited number of labels available from each user and mitigates the effect of label noise. At the same time, maintaining a manageable number of reward models reduces the validation burden at deployment. Experiments across diverse simulated locomotion environments show that PREC groups users who label different trajectory subsets into preference-coherent clusters more accurately than baseline methods. Under sparse and noisy feedback, policies trained with PREC improve all three social welfare metrics over an existing single shared-policy user-alignment approach and even outperform per-user alignment approaches.
Comments23 pages, 20 figures