arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

声音作为句柄:利用冻结文本大语言模型推理说话人身份

Voices as Handles: Reasoning about Speaker Identity with Frozen Text LLMs

Runqiu Xu, Zhisheng Zheng, David Harwath

arXiv 2609.38501首次发表:更新:

发表机构

The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出说话人句柄(Speaker Handles),通过轻量投影器将声学身份映射为软令牌,使冻结文本大语言模型能跨会话推理说话人归属,在VoxCeleb1和SpeakerBind上分别达到97.40-98.36%和70.40%的准确率。

AI 中文摘要

多用户语音代理必须在对话会话中跟踪谁说了什么。文本大语言模型是此类代理的理想骨干,但仅凭转录文本无法暴露声学说话人身份,使得模型缺乏跨会话将信息与说话人关联的持久参考。我们通过引入说话人句柄(Speaker Handles)来解决这一空白,这是一种软令牌表示,可将声学说话人身份暴露给冻结的文本大语言模型,以支持跨会话的说话人相关推理。一个三阶段课程训练了一个轻量级投影器,其参数少于骨干参数的0.1%,用于将说话人嵌入映射到这些句柄中。由于现有基准中文本线索可能部分揭示事实归属,因此验证生成的句柄是否真正支持跨会话说话人相关推理具有挑战性。为此,我们提出了SpeakerBind,一个受控的共享代理基准,其中用户间的重叠事实要求正确的跨会话说话人归属。说话人句柄在VoxCeleb1上达到97.40-98.36%的准确率,在SpeakerBind上达到70.40%,接近71.88%的上限。这些结果表明,所提出的说话人句柄为将声学说话人身份集成到冻结文本大语言模型中进行说话人内容推理提供了一种高效途径。

英文摘要

Multi-user voice agents must track who said what across dialogue sessions. Text LLMs are attractive backbones for such agents, but transcripts alone do not expose acoustic speaker identity, leaving the model without a persistent reference for linking information to speakers across sessions. We address this gap by introducing Speaker Handles, soft-token representations that expose acoustic speaker identity to a frozen text LLM for cross-session speaker-dependent reasoning. A three-stage curriculum trains a lightweight projector, with fewer than 0.1% of the backbone's parameters, to map speaker embeddings into these handles. Establishing whether the resulting handles truly support cross-session speaker-dependent reasoning is challenging with existing benchmarks because textual cues can partially reveal fact ownership. We therefore present SpeakerBind, a controlled shared-agent benchmark in which overlapping facts across users require correct cross-session speaker attribution. Speaker Handles achieve 97.40-98.36% accuracy on VoxCeleb1 and 70.40% on SpeakerBind, close to the 71.88% topline. These results show that the proposed Speaker Handles provide an efficient way to integrate acoustic speaker identity into frozen text LLMs for speaker-content reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑