arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

INSPIRE:指令感知语音检索基准

INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval

Chen-An Li, Hung-yi Lee

arXiv 2608.16203首次发表:更新:

发表机构

National Taiwan University; NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)(台湾大学; 台湾大学人工智能卓越研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有语音检索系统无法适配多样用户意图的问题,构建首个指令感知语音检索基准INSPIRE,评估四类检索范式,发现当前无方法能稳健处理所有检索意图,凸显统一架构的必要性。

AI 中文摘要

现有语音检索系统依赖固定的相似度匹配,无法适配多样的用户意图。我们推出首个指令感知语音检索基准INSPIRE,其中自然语言指令可动态指定相关性标准,包括语义内容、说话人身份、说话风格、环境声音及其组合。我们评估了四种检索范式:大型音频-语言模型、级联流水线、自监督语音模型、对比音频-语言模型。结果显示,当前没有任何方法能稳健处理所有检索意图:基于文本的方法在语义检索上表现相对较好,但在副语言属性上存在困难;而基于语音的模型在捕捉声学特性方面表现稍好,但在遵循指令时表现不佳。这些发现凸显了对具备指令感知语音检索能力的统一架构的需求。

英文摘要

Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language instructions dynamically specify relevance criteria, including semantic content, speaker identity, speaking style, environmental sounds, and their combinations. We evaluate four retrieval paradigms: large audio-language models, cascaded pipelines, self-supervised speech models, and contrastive audio-language models. Our results reveal that no current method robustly handles all retrieval intents. Text-based approaches perform relatively better at semantic retrieval but struggle with paralinguistic attributes, while speech-based models are moderately better at capturing acoustic properties but falter at following instructions. These findings highlight the need for unified architectures capable of instruction-aware speech retrieval.

CommentsInterspeech 2026 long paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑