arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用错误示例向人类传授机器人策略

Teaching Robot Policies to Humans Using Erroneous Examples

Rithika Narayan, Suresh Kumaar Jayaraman, Henny Admoni

arXiv 2608.29023首次发表:更新:

AI 中文总结

本研究提出用错误示例扩展现有框架,通过用户研究发现该方法可提升人类对机器人策略的长期记忆,类逆强化学习推理者表现最优,以推进机器人教学人类策略的方法。

AI 中文摘要

人机协作描述了人类与自主智能体共同完成目标的过程,该过程在机器人策略(即不同情境下的行为)对人类透明时效果最佳。基于演示的解释一直是人机协作研究的重点,该领域常借鉴教育领域的文献以改进向人类传授机器人策略的方式。然而,尚无单一教学方法在不同领域、难度、学习者及其他变量中被证明有效;如何最有效地向人类传授机器人策略仍是未解决的问题。在传统课堂中,学习者会接触错误示例,通过反思并纠正错误答案来理解学习概念时的常见误区。我们提出使用错误示例来传授机器人策略,扩展了现有的策略教学框架。我们开展了一项用户研究,参与者观看机器人行为的错误演示并纠正动作以匹配实际策略。研究结果表明,观看这些错误演示并在预测机器人动作时表述自身推理,可提升对策略的长期记忆,与课堂中错误示例的效果一致。我们还将参与者分为不同学习风格,发现使用类逆强化学习推理的参与者在策略预测任务中表现最佳。本研究旨在推进机器人向人类传授其策略的方法。

英文摘要

Human-robot collaboration describes the process of humans and autonomous agents working together to accomplish common goals. This process is facilitated best when robot policies, or behaviors in different situations, are made transparent to humans. Demonstration-based explanations have been a focus of human-robot collaboration research, and the field has frequently drawn upon literature from education to improve how humans are taught robot policies. However, no single teaching method has been proven effective across domains, difficulties, learners, and other variables; the question of how humans can most effectively be taught robot policies remains open. In traditional classrooms, learners are shown erroneous examples, in which they reflect on and correct incorrect responses to understand common pitfalls when learning a concept. We propose using erroneous examples to teach robot policies, extending an existing policy teaching framework. We conduct a user study in which participants view incorrect demonstrations of robot behavior and correct the actions to align with the actual policy. Our findings suggest that viewing these incorrect demonstrations and verbalizing one's reasoning in predicting a robot's actions improves retention of the policy over time, in agreement with the effect of erroneous examples in classrooms. We also categorize participants into distinct learning styles and establish that participants using inverse reinforcement learning-like reasoning perform best on policy prediction tasks. With this work, we aim to advance the methods by which robots educate humans on their policies.

Comments31 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑