发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对唤醒词在协同AR中的误激活与对话对象识别难题,提出基于注视调用、具身可见状态与回合管理的交互设计,经25人实验验证,98.8%的语音尝试在语音前完成注视获取,并揭示释放阶段状态传达问题。
AI 中文摘要
唤醒词面临固有的权衡:更高的灵敏度会减少漏掉的命令,但会增加意外激活。在协同定位的增强现实(AR)中,系统还必须确定用户是在与助手交谈还是与附近的人交谈。我们引入了对这一问题的全新视角:将基于注视的寻址与具身助手、可见的聆听状态以及跨激活、持续交互和释放的回合管理相结合。在我们的实现中,用户看向场景中锚定的助手,并在其身体上看到激活进度和聆听状态。持续注视打开局部交互通道;当注意力回到任务时,语音和播放保持通道开启;不活动则关闭通道。我们与25名参与者评估了该设计。在选定助手的放置位置和驻留时长后,每位参与者与搭档合作规划旅行,同时使用助手。研究记录了500个交互结果,包括496次面向助手的语音尝试。其中490次尝试(98.8%)在语音之前完成了注视获取:467次在通道保持开启时开始,而23次在通道释放后开始。其余6次尝试在获取完成前开始。在102次审查的尝试中,注视在获取后、语音前离开,配置的保留策略为79次保持通道开启,23次在之前释放。结果表明,具有可见状态的具身目标可以支持清晰地进入助手交互,同时揭示了释放时的不同问题:用户将视线转回工作后,可能看不到助手已停止聆听。我们贡献了已实现的注视调用设计以及关于在视觉注意力转移到别处后传达助手状态的设计启示。
英文摘要
Wake words face an inherent trade-off: higher sensitivity reduces missed commands but increases accidental activations. In co-located augmented reality (AR), the system must also determine whether the user is speaking to the assistant or to a nearby person. We introduce a new perspective on this problem: combining gaze-based address with an embodied assistant, visible listening status, and turn management across activation, continued interaction, and release. In our implementation, users look at an assistant anchored in the scene and see activation progress and listening status on its body. Sustained gaze opens a local interaction channel; speech and playback keep it open when attention returns to the task; and inactivity closes it. We evaluated this design with 25 participants. After selecting the assistant's placement and dwell duration, each participant worked with a partner to plan a trip while using the assistant. The study recorded 500 interaction outcomes, including 496 assistant-directed utterance attempts. Gaze acquisition completed before speech for 490 of these attempts (98.8\%): 467 began while the channel remained open, whereas 23 began after it had been released. The other 6 attempts began before acquisition completed. Among 102 reviewed attempts in which gaze left after acquisition but before speech, the configured retention policy kept the channel open for 79 and released it before 23. The results show that an embodied target with visible status can support clear entry into an assistant interaction while revealing a different problem at release: after users look back to their work, they may not see that the assistant has stopped listening. We contribute the implemented gaze-invocation design and design implications for communicating assistant state after visual attention moves elsewhere.
Comments11 pages, 9 figures, 2 tables