arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于神经与嵌入的说话人日志的无训练亲和融合

Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization

Yehoshua Dissen, Joseph Keshet, Eduard Golshtein

arXiv 2609.39162首次发表:更新:

发表机构

Linguana; Technion – Israel Institute of Technology(Linguana; 以色列理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出无训练亲和融合(TFAF),将神经日志器的说话人分区与嵌入亲和矩阵结合,无需额外训练,在AMI和CALLHOME上提升日志性能。

AI 中文摘要

基于说话人嵌入和神经日志的说话人日志系统利用了互补的说话人信息形式,但它们的中间表示并不直接兼容。我们引入了无训练亲和融合(TFAF),该方法将神经日志器推断出的说话人结构整合到基于嵌入的日志系统中。神经说话人分区用于条件化局部说话人表示,我们从中构建一个连续的亲和矩阵,并在单一全局聚类步骤之前将其与基于嵌入的声学亲和相结合。该方法无需额外训练、共享嵌入空间、说话人标签对齐或神经日志器说话人数的硬性转移。在AMI和CALLHOME上的实验显示,与两个组成系统相比,DER持续改进;在AMI上,融合还改善了说话人属性转录。消融研究表明,神经说话人分区贡献了大部分增益,而保留连续的基于嵌入的亲和比硬分区融合提供了额外的好处。

英文摘要

Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermediate representations are not directly compatible. We introduce Training-Free Affinity Fusion (TFAF), which integrates the speaker structure inferred by a neural diarizer into an embedding-based diarization system. The neural speaker partition is used to condition local speaker representations, from which we construct a continuous affinity matrix and combine it with the embedding-based acoustic affinity before a single global clustering step. The method requires no additional training, shared embedding space, speaker-label alignment, or hard transfer of the neural diarizer's speaker count. Experiments on AMI and CALLHOME show consistent DER improvements over both constituent systems; on AMI, fusion also improves speaker-attributed transcription. Ablations show that the neural speaker partition accounts for most of the gain, while retaining the continuous embedding-based affinities provides additional benefit over hard partition fusion.

Commentssubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑