从 Wizard-of-Oz 人机对话收集到机器人响应决策分类法:辅助驾驶交互的回顾性分析
From Wizard-of-Oz Human-Robot Dialogue Collection to a Taxonomy of Robot Response Decisions: A Retrospective Analysis of Assistive Pilot Interactions
- Saint Louis University(圣路易斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过回顾性分析Wizard-of-Oz辅助驾驶交互数据,提出包含六种响应模式和四种歧义类型的层次分类法,并验证了其标注可靠性及用于训练视觉语言模型的可行性。
AI中文摘要:
在日常生活室内环境中遵循自然语言指令的机器人必须处理不完整的人类话语。指令常常省略关键信息,例如视野外物体的身份、预期目的地或用户的目标。现有数据集包含很少的真实情境对话,并且很少提供基于实践的准则来决定机器人何时应行动、确认、澄清或拒绝。我们回顾性分析了一项试点 Wizard-of-Oz 研究,其中五名参与者使用安装在轮椅上的移动机械臂执行日常室内任务,包括开门、开抽屉、喂食、饮水和清洁,而向导在无正式通信策略的情况下进行响应。这保留了真实的用户行为,但产生了不一致的机器人侧决策,促使我们提出一个明确的决策方案。从40个片段中,我们推导出一个包含六种响应模式(ANSWER、REPORT_DONE、REFUSE、CONFIRM、CLARIFY、ACT)和四种歧义类型(意图、指称、空间、可理解性)的层次分类法。两名人类标注者和一名AI标注者将该方案应用于试点数据。干净标签率分别为91%和89%,在决策点、模式和歧义级别上,人类-人类和人类-AI比较的Cohen's kappa范围从0.72到0.95。在ACT和CLARIFY上对LLaVA-1.6-7B进行基于分类法标签的微调,表明使用我们分类法的注释训练视觉语言模型的可行性。决策点识别和REPORT_DONE中剩余的边界案例促使我们提出一个受约束的协议,以实现更一致的对话收集。
英文摘要:
Robots that follow natural-language instructions in everyday indoor environments must act on incomplete human utterances. Instructions often omit essential information, such as the identity of an out-of-view object, an intended destination, or the user's goal. Existing datasets contain little real-world situated dialogue and provide few practice-grounded criteria for deciding when a robot should act, confirm, clarify, or refuse. We retrospectively analyze a pilot Wizard-of-Oz study in which five participants performed everyday indoor tasks, including door opening, drawer opening, feeding, drinking, and cleaning, with a wheelchair-mounted mobile manipulator while the wizard responded without a formal communication policy. This preserved authentic user behavior but produced inconsistent robot-side decisions, motivating an explicit decision scheme. From 40 episodes, we derived a hierarchical taxonomy of six response modes (ANSWER, REPORT_DONE, REFUSE, CONFIRM, CLARIFY, ACT) and four ambiguity types (intent, referential, spatial, intelligibility). Two human annotators and an AI annotator applied the scheme to the pilot data. Clean-label rates were 91% and 89%, and Cohen's ranged from 0.72 to 0.95 across decision-point, mode, and ambiguity levels for both human-human and human-AI comparisons. Fine-tuning LLaVA-1.6-7B on taxonomy-derived labels for ACT and CLARIFY indicates the feasibility of training vision-language models using annotations from our taxonomy. Remaining boundary cases in decision-point identification and REPORT_DONE motivate a constrained protocol for more consistent dialogue collection.