AI 中文总结
Voice-Light提出级联全双工语音智能体,结合因果轮换与投机式生成,实现低延迟交互,并通过混合控制器在真实场景中平衡误切断与轮次召回率。
AI 中文摘要
自然的口语交互需要的不仅仅是流式语音识别、语言生成和语音合成:系统必须在不因每次确认而取消的情况下对重叠语音做出反应,在轮次确定之前准备响应,并确保被取消的音频不会进入对话历史。我们提出了Voice-Light,一个级联全双工语音智能体,它结合了即时声学起始、与流式语音识别编码器共享的因果适配器、可逆播放控制以及私有投机式响应生成。结构化工具调用与可听见的桥接语音并发执行,而浏览器确认使渲染的音频对持久历史具有权威性。在1,673个真实对话静音候选上的锁定评估发现,较早学习的完成检查点保持了2.70%的误切断率,但端到端轮次召回率仅为12.53%,而Silero计时策略的召回率为95.60%。因此,部署的系统保留了混合控制器,而不是声称用学习策略替代。在三次无脚本的操作员运行的麦克风会话中,36个测量的响应轮次从最终VAD端点到首个服务器音频的中位数为758毫秒;21个轮次低于800毫秒。这些会话是仪器化的案例研究,而非受控的用户评估。我们发布了支持该结果的合成数据、模型工件、评估代码和摘要、源代码以及部署配置。
英文摘要
Natural spoken interaction requires more than streaming ASR, language generation, and speech synthesis: a system must react to overlap without canceling on every acknowledgment, prepare a response before a turn is certain, and ensure canceled audio cannot enter conversation history. We present Voice-Light, a full-duplex cascaded voice agent that combines immediate acoustic onset, a causal adapter sharing a streaming ASR encoder, reversible playback control, and private speculative response generation. Structured tool calls execute concurrently with audible bridge speech, while browser acknowledgments make rendered audio authoritative for durable history. Locked evaluation on 1,673 real-conversation silence candidates found that an earlier learned completion checkpoint preserved a 2.70% false-cutoff rate but reached only 12.53% end-of-turn recall, compared with 95.60% for a Silero timing policy. The deployed system therefore retains a hybrid controller rather than claiming a learned-policy replacement. Across three unscripted operator-run microphone sessions, 36 measured response turns had a 758 ms median from final VAD endpoint to first server audio; 21 turns were below 800 ms. These sessions are an instrumented case study, not a controlled user evaluation. We release the synthetic data, model artifacts, evaluation code and summaries, source code, and deployment configuration supporting the result.
Comments9 pages, 4 figures, 6 tables. Code, datasets, and model artifacts: https://github.com/BertilBraun/Voice-Light ; live demo: https://voice.bertil-braun.de