发表机构
TCL Research Europe; University of Warsaw(TCL欧洲研究院; 华沙大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种扩展传统关键词检测的上下文触发唤醒系统,通过推理区分用户命令与无关语音,并生成62.3小时可控多说话人对话语料库验证其有效性。
AI 中文摘要
唤醒词检测是虚拟助手的关键组成部分,是无缝用户交互的门户。本文介绍了一种新颖的唤醒系统,将传统的直接关键词检测扩展为上下文触发检测。在初始唤醒词激活后,系统利用推理来区分用户命令与无关语音,确保高效且上下文感知的交互。我们提出了一种数据生成架构,生成了包含直接调用、上下文后续和非目标语音的62.3小时可控多说话人对话语料库。实验结果表明,所提方法在多种合成对话场景中的有效性。我们发布了代码、数据集和训练模型,以促进可复现性和智能助手技术的进一步进步。
英文摘要
Wake word detection is a critical component of virtual assistants, serving as the gateway to seamless user interactions. This paper introduces a novel wake-up system that extends traditional direct keyword detection with contextual trigger detection. After an initial wake word activation, the system uses reasoning to distinguish between user commands and unrelated speech, ensuring efficient and context-aware engagement. We present a data generation architecture that produces a 62.3-hour corpus of controllable multi-speaker conversations containing direct invocations, contextual follow-ups, and non-addressed speech. Experimental results demonstrate the effectiveness of the proposed approach across diverse synthetic conversational scenarios. We release the code, dataset and trained models to promote reproducibility and further advancements in intelligent assistant technologies.