arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考人类反馈对人机协作中偏好学习的影响

Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Collaboration

Qiping Zhang, Kate Candon, Debasmita Ghose, Marynel Vázquez

arXiv 2609.13982首次发表:更新:

发表机构

Yale University(耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对人机协作中偏好学习,提出IMPLIED方法,通过建模人类反馈的蕴含关系,改进奖励模型学习,提高行为适应性。

AI 中文摘要

在人机交互中,学习代表人类对机器人行为偏好的奖励模型的标准方法包括三个步骤。首先,机器人从人类反馈(例如,正面或负面的二元反馈)中收集有限的直接证据。然后,机器人利用直接证据,通过固定的蕴含规则,为可行但未被选择的行为推导出接受或拒绝的标签。最后,机器人使用直接证据和推导出的证据更新奖励模型。不幸的是,固定规则会阻碍偏好学习:在一项包含两个协作模拟环境的用户研究中,人类提供的蕴含标签通常与标准固定规则不同,并且使用人类标签显著改善了使用隐式和显式反馈偏好学习(PIE)算法的奖励学习。因此,我们提出了IMPLIED,一种蕴含建模方法,该方法将固定规则的蕴含作为初始指导,同时学习随时间推断和修订接受和拒绝的行为标签。在记录的人机交互轨迹和物理机器人制作比萨研究的评估中,IMPLIED预测人类蕴含的准确性高于固定规则方法和LLM基线,接近人类标签预言机的性能。反过来,IMPLIED减少了偏好估计误差,并使机器人行为相对于基线更常与组合奖励(包括真实偏好奖励和任务特定奖励)保持一致。通过学习推理人类反馈的蕴含,这项工作能够在人机协作过程中实现更忠实、更高效的机器人行为适应。

英文摘要

In Human-Robot Interaction, the standard approach to learn a reward model that represents human preferences for robot behavior consists of three steps. First, the robot collects limited direct evidence from human feedback (e.g., positive or negative binary feedback). Then, the robot utilizes the direct evidence to derive accepted or rejected labels to feasible but unchosen actions using fixed implication rules. Finally, the robot updates the reward model with both the direct and derived evidence. Unfortunately, the fixed rule can hinder preference learning: in a user study with two collaborative simulation environments, human-provided implication labels often differed from the standard fixed rule, and using the human labels substantially improved reward learning with the Preference Learning from Implicit and Explicit Feedback (PIE) algorithm. Consequently, we propose IMPLIED, an implication modeling method that treats fixed-rule implications as an initial guide while learning to infer and revise accepted and rejected action labels over time. Across evaluations on recorded human-robot interaction trajectories and a physical robot pizza-making study, IMPLIED predicts human implications more accurately than the fixed rule approach and LLM baselines, approaching the performance of a human-label oracle. In turn, IMPLIED reduces preference-estimation error and leads to robot actions that are more often rational with respect to a combined reward (which includes the true preference reward and a task-specific reward) compared to baselines. By learning to reason about the implications of human feedback, this work enables more faithful and efficient robot behavior adaptation during human-robot collaboration.

CommentsConference on Robot Learning (CoRL) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑