arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TS-SP:在音频大语言模型中学习说话人保持表示

TS-SP: Learning Speaker-Preserving Representations in Audio Large Language Models

Junjie Li, Zheng Liang, Zhe Li, Tianchi Liu, Kong Aik Lee

arXiv 2610.04887首次发表:更新:

发表机构

The Hong Kong Polytechnic University; The University of Hong Kong; National University of Singapore(香港理工大学; 香港大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出TS-SP两阶段适配框架,通过LoRA在Qwen2.5-Omni-7B上学习说话人保持表示,显著降低Vox1-O等错误率至4.37%,支持跨域验证。

AI 中文摘要

音频大语言模型(ALLM)能够理解语音内容,但其利用说话人身份进行验证的能力仍然有限。我们提出TS-SP(两阶段说话人保持),一种参数高效的框架,用于学习说话人保持表示并使其可被ALLM的语言模型组件访问。我们在Qwen2.5-Omni-7B上实例化并评估TS-SP。首先,我们使用说话人身份监督来适配音频编码器。然后,我们冻结适配后的编码器,并训练语言模型以比较说话人。两个阶段均使用低秩适配(LoRA),保持预训练基础权重固定。在Vox1-O上,TS-SP将等错误率(EER)从配对损失适配基线的7.01%降低到4.37%。在未见提示下,EER保持在4.31%至4.79%之间。在CN-Celeb上的跨域评估产生了与基线相似的EER,但在原生决策阈值下准确率较低。这些发现支持两阶段适配以改善所评估骨干上的说话人验证。

英文摘要

Audio large language models (ALLMs) can understand speech content, yet their ability to use speaker identity for verification remains limited. We propose TS-SP (Two-Stage Speaker Preservation), a parameter-efficient framework for learning speaker-preserving representations and making them accessible to an ALLM's language-model component. We instantiate and evaluate TS-SP on Qwen2.5-Omni-7B. First, we adapt the audio encoder with speaker identity supervision. We then freeze the adapted encoder and train the language model to compare speakers. Both stages use low-rank adaptation (LoRA), keeping the pretrained base weights fixed. On Vox1-O, TS-SP reduces the equal error rate (EER) from 7.01\% for the Paired Loss Adaptation Baseline to 4.37\%. EER remains within 4.31--4.79\% under unseen prompts. Cross-domain evaluation on CN-Celeb yields a similar EER to the baseline, but lower accuracy at the native decision threshold. These findings support two-stage adaptation for improving speaker verification on the evaluated backbone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑