AI 中文总结
该研究针对奖励无法直接指定的场景,提出人机团队中教师与学习者通过二阶心理理论同步信念的方法,仿真显示该方法能提升偏好学习性能并修复模型漂移问题。
AI 中文摘要
当无法直接指定奖励时,比较反馈(即询问人们两种行为中更偏好哪一种)已成为使机器人与智能体行为契合人类意图的标准方式。基于偏好的奖励学习通常将人类教师视为回答学习者生成查询的被动预言家。我们认为这会丧失教师的核心优势:对目标的认知。了解目标的教师能比任何学习者驱动的获取策略更高效地构建训练样本,且该优势会随奖励的特征维度增大而扩大。然而,利用此优势需要对学习者当前的知识有准确模型。因此,我们将偏好学习重新定义为耦合两个行为模型的人机自主团队问题:教师维护学习者模型以设计信息丰富的课程,学习者维护教师模型的二阶模型,生成结构化偏好约束(理解语句)以保持教师的学习者模型同步。在仿真中,知情教师的表现优于学习者主导的选择;教师在交替教师下的模型漂移会削弱该优势;而理解语句可修复此问题,当教师对学习者的误差集中在特定方向而非均匀分布时,二阶心理理论(ToM-2)语句的表现优于平均信念语句。
英文摘要
Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly. Preference-based reward learning typically casts the human teacher as a passive oracle answering learner-generated queries. We argue this forfeits the teacher's defining advantage: knowledge of the objective. A teacher who knows the target can construct training examples more efficiently than any learner-driven acquisition strategy, an advantage that widens as the reward's feature dimension grows. However, exploiting this advantage requires an accurate model of what the learner currently knows. We therefore recast preference learning as a human-autonomy team problem coupling two behavioral models: the teacher maintains a model of the learner to design an informative curriculum, and the learner maintains a second-order model of the teacher's model, emitting structured preference constraints (understanding statements) that keep the teacher's model of the learner synchronized. In simulation, an informed teacher outperforms learner-led selection; teacher-model drift under alternating teachers erodes this advantage; and understanding statements repair it, with second-order (ToM-2) statements outperforming mean-belief statements when the teacher's error about the learner is concentrated in a particular direction rather than spread evenly.