arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33212cs.CLcs.SDeess.AS

CoLMbo-SV:一种用于可解释说话人验证的接地语言模型

CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification

Massa Baali, Sarthak Bisht, Ziyue Qiu, Joseph Konan, Rita Singh, Bhiksha Raj

首次发表
浏览论文内容

中文总结 AI 辅助

CoLMbo-SV通过将预训练说话人编码器与语言模型结合,并引入VoxReason数据集和评估框架,在VoxCeleb1-O上实现0.99%等错误率,提供可检查的声学比较报告,显著提升可解释说话人验证的准确性与解释质量。

中文摘要 AI 辅助

说话人验证系统实现了高准确率,但对其判断背后的声学证据提供的说明甚少。要使这些系统可检查,需要暴露可解释的证据,同时保留其决策所依赖的更丰富信息。我们提出了\ extbf{CoLMbo-SV},一种说话人语言模型,它将强大的说话人区分能力与结构化、声学接地的比较报告相结合。通过将预训练的说话人编码器连接到语言模型并提供明确的声学测量,CoLMbo-SV使语音比较可检查,而无需将验证限制在其报告中口头化的证据上。我们还引入了\ extbf{VoxReason},即带有测量声学属性和经过数值与定性检查过滤的比较报告的配对录音,为这种组合能力提供监督。我们还开发了一个评估框架,该框架区分了说话人表示编码的声学信息、影响验证分数的因素以及生成的报告所讨论的内容。在VoxCeleb1-O上,CoLMbo-SV实现了0.99%的等错误率,相对于在VoxReason上微调的最强音频语言基线,验证错误减少了约80%,同时达到了0.82的数值接地分数。我们的分析进一步表明,声学正确性和决策相关性是解释的不同属性,揭示了数值接地指标所遗漏的差距。总之,这些贡献显著推进了音频语言说话人验证,使其准确率接近专用说话人编码器,同时增加了可检查的声学报告,并建立了一个将自然语言解释与其所解释的决策联系起来的实证框架。

英文摘要

Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-SV}, a speaker language model that combines strong speaker discrimination with structured, acoustically grounded comparison reports. By connecting a pretrained speaker encoder to a language model and supplying explicit acoustic measurements, CoLMbo-SV makes voice comparisons inspectable without restricting verification to the evidence verbalized in its reports. We additionally introduce \textbf{VoxReason}, paired recordings with measured acoustic properties and comparison reports filtered through numerical and qualitative checks, providing supervision for this combined capability. We also develop an evaluation framework that separates what acoustic information a speaker representation encodes, what influences the verification score, and what the generated report discusses. On VoxCeleb1-O, CoLMbo-SV achieves 0.99\% EER, reducing verification error by approximately 80\% relative to the strongest audio-language baseline fine-tuned on VoxReason, while attaining a numerical-grounding score of 0.82. Our analysis further demonstrates that acoustic correctness and decision relevance are distinct properties of an explanation, exposing a gap that numerical-grounding metrics miss. Together, these contributions substantially advance audio-language speaker verification, bring its accuracy toward that of dedicated speaker encoders while adding checkable acoustic reporting, and establish an empirical framework for connecting natural-language explanations to the decisions they explain.

↑