前景语音活动检测:从监督中学习说话人选择性
Foreground Voice Activity Detection: Learning Speaker Selectivity from Supervision
- Recho Inc.(Recho公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出前景语音活动检测(FVAD),通过自动生成的竞争说话人混合增强训练,使轻量级流式模型Mamba-FVAD在保持传统VAD性能的同时,显著提升前景说话人选择性,优于商业系统。
AI中文摘要:
语音活动检测(VAD)是大多数语音助手流水线的前端,然而生产环境中的检测器将所有人类语音(包括背景说话人)都视为有效活动;在拥挤场景中,这会导致识别过载、打断轮转,并触发虚假的打断插入。我们将前景语音活动检测(FVAD)形式化:这是一种帧同步、无需注册的任务,其中只有主导说话人(由持续存在而非瞬时响度定义)被视为正样本,并且当仅存在单个说话人时,该任务退化为传统VAD。我们表明,前景选择性主要由训练监督决定:关键要素是一种数据增强方案,将仅前景标签与竞争说话人混合配对,完全自动生成,无需人工标注。为了量化选择性,我们引入了背景误报率(BG-FAR),以前景F1为门控,并构建了一个受控基准Mix-Interference,辅以适配的VOiCES用于真实远场评估。在同等规模的骨干网络下,Mamba和LSTM表现相当,而更长上下文的注意力模型并未更好,这表明训练监督在实现前景选择性方面比时间建模能力发挥的作用大得多。由此产生的轻量级流式模型Mamba-FVAD在前景选择性上优于商业VAD和基于注册的说话人感知系统,同时在传统VAD上保持竞争力,每帧CPU延迟为1-2毫秒。
英文摘要:
Voice activity detection (VAD) fronts most voice-agent pipelines, yet production detectors treat all human speech, background talkers included, as valid activity; in crowded settings this floods recognition, stalls turn-taking, and triggers false barge-in. We formalize Foreground VAD (FVAD): a frame-synchronous, enrollment-free task in which only the dominant speaker, defined by sustained presence rather than instantaneous loudness, is positive, and which reduces to conventional VAD when a single speaker is present. We show that foreground selectivity is largely governed by training supervision: the crucial ingredient is an augmentation recipe pairing foreground-only labels with competing-speaker mixing, generated fully automatically without human annotation. To quantify selectivity we introduce the Background False-Alarm Rate (BG-FAR), gated by foreground F1, and build a controlled benchmark, Mix-Interference, complemented by an adapted VOiCES for real-world far-field evaluation. Across equal-size backbones, Mamba and LSTM perform on par while a longer-context attention model is no better, suggesting that training supervision plays a substantially larger role than temporal modeling capacity in achieving foreground selectivity. The resulting lightweight streaming model, Mamba-FVAD, outperforms commercial VADs and enrollment-based speaker-aware systems in foreground selectivity while staying competitive on conventional VAD, at 1-2 ms per-frame CPU latency.