AI 中文总结
DUPAR通过双路径对话检索框架,利用跨轮证据缓存和语音编码器,在保持检索准确率的同时实现3.75倍加速,并显著提升噪声环境下的检索性能。
AI 中文摘要
基于外部知识的语音助手通常使用自动语音识别(ASR)将语音查询转录为文本,然后从文本知识库中检索证据。这种级联方式增加了延迟并传播识别错误,而直接进行语音检索则容易受到跨模态错位的影响。为解决这些局限,我们提出了DUPAR,一种具有互补的慢速和快速路径的对话检索框架。快速路径使用与冻结的BGE-M3文本嵌入对齐的任务自适应音频编码器,以搜索跨轮证据缓存。当缓存置信度不足时,慢速路径融合使用音频和ASR转录嵌入的全索引检索,所选证据通过一跳图扩展刷新下一轮的证据缓存。在特定领域知识库上,我们训练的音频编码器在干净语音上接近文本检索的准确率,查询端速度比ASR+文本编码器快3.75倍。在噪声基准上将平均Recall@10从0.771提升至0.875,并在合成说话风格上将整体Recall@1提高了4.2个百分点。与全索引音频检索相比,当上一轮检索到正确证据且后续目标为一跳相邻块时,跨轮证据缓存显著减少了检索错误。
英文摘要
Voice assistants grounded in external knowledge typically use automatic speech recognition (ASR) to transcribe speech queries before retrieving evidence from textual knowledge bases. This cascade adds latency and propagates recognition errors, whereas direct speech retrieval is vulnerable to cross-modal misalignment. To address these limitations, we propose DUPAR, a conversational retrieval framework with complementary slow and fast paths. The fast path uses a task-adapted audio encoder aligned with frozen BGE-M3 text embeddings to search a cross-turn evidence cache. When cache confidence is insufficient, the slow path fuses full-index retrieval using audio and ASR-transcript embeddings, and the selected evidence refreshes the next-turn evidence cache through one-hop graph expansion. On a domain-specific knowledge base, our trained audio encoder approaches text-retrieval accuracy on clean speech with a 3.75$\times$ query-side speedup over ASR + Text Encoder. It raises average Recall@10 from 0.771 to 0.875 on the noise benchmark and improves overall Recall@1 by 4.2 percentage points across synthesized speaking styles. Compared with full-index audio retrieval, cross-turn evidence caching significantly reduces retrieval errors when the previous turn retrieves correct evidence and the follow-up targets a one-hop neighboring chunk.
Comments5 pages, 4 figures