AI 中文总结
针对全双工语音智能体,IndicFDB将基准扩展到印度十种语言,提供12,350个样本,并解决事件挖掘、时序评估和跨语言评判三大挑战,揭示商业API与开放模型间的性能权衡。
AI 中文摘要
全双工语音智能体必须实时处理停顿、轮流发言、回馈以及响应用户的打断。Full-Duplex-Bench评估了这些行为,但其仅含英语的语料库以及对带词级时间戳的自动语音识别(ASR)和英语提示的大语言模型(LLM)评判器的依赖,使其难以扩展到印度语言。我们提出了IndicFDB,将其扩展到印度使用的十种语言,包含12,350个样本,几乎是原始样本数的17倍。我们解决了三个挑战:在多语言语音中寻找对话事件、在缺乏可靠的词级对齐的情况下评估其时间性,以及跨语言评判响应。我们利用语音活动检测(VAD)从大约50,000小时的通道分离对话中挖掘停顿处理、轮流发言和回馈样本,并构建了人工验证的合成用户打断样本。语言无关的VAD启发式方法用于评估时间性,而一个开放权重的转录和翻译流水线将响应转换为英语,以供大语言模型(LLM)对相关性和质量进行评分。在七个语音智能体中,商业API在不同语言间表现出出乎意料的一致性行为,但要么速度快,要么对停顿稳健,从未两者兼得,而单语开放全双工模型则进一步暴露了回馈、响应质量和延迟之间的权衡。
英文摘要
Full-duplex voice agents must handle pauses, take turns, backchannel, and respond to user interruptions in real time. Full-Duplex-Bench evaluates these behaviors, but its English-only corpus and reliance on word-timestamped ASR and an English-prompted LLM judge make it difficult to extend to Indian languages. We introduce IndicFDB, which extends it to ten languages spoken in India with 12,350 samples, nearly 17 times as many as the original. We address three challenges: finding conversational events in multilingual speech, evaluating their timing without reliable word-level alignment, and judging responses across languages. We mine pause handling, turn taking, and backchanneling samples from roughly 50,000 hours of channel-separated conversations using voice activity detection (VAD), and construct human-validated synthetic user interruption samples. Language-independent VAD heuristics evaluate timing, while an open-weight transcription and translation pipeline converts responses to English for LLM ratings of relevance and quality. Across seven voice agents, commercial APIs show unexpectedly consistent behavior across languages but are either fast or robust to pauses, never both, while monolingual open full-duplex models expose further tradeoffs among backchanneling, response quality, and latency.