arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语码转换口语语言识别作为多标签集合预测

Code-Switching Spoken Language Identification as Multi-Label Set Prediction

Shunsuke Mitsumori, Matthew Wiesner, Shigeo Morishima, Shinji Watanabe

arXiv 2610.01450首次发表:更新:

发表机构

Waseda University; Johns Hopkins University; Carnegie Mellon University(早稻田大学; 约翰斯·霍普金斯大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出将语码转换语音的语言识别建模为多标签集合预测,通过集合生成器直接输出话语语言,分析揭示了oracle基数、阈值不稳定等关键挑战。

AI 中文摘要

语码转换(CS)语音会泄漏通过用于筛选大规模语音语料库的单语语言识别(LID)过滤器,因此需要具有CS感知的语言识别(CS-LID)。我们将话语级CS-LID表述为多标签语言集合预测,并提出一个集合生成器,直接输出话语中的语言,并将其与原子对和基于分数的分类基线进行比较。Oracle Top-k是最强的基线,但阈值化失败,因为没有单一阈值可以将CS与单语语音分开。我们的集合生成器在未假设语言数量的情况下,对未见过的语言对预测正确的语言数量,但在精确集合准确率上不如Oracle Top-k。我们的分析确定了稳健CS-LID的关键障碍:oracle基数、阈值不稳定性、CS训练数据中的语言偏差以及合成到真实的差距。

英文摘要

Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utterance, comparing it against atomic-pair and score-based classification baselines. Oracle Top-k is the strongest baseline, but thresholding fails because no single threshold separates CS from monolingual speech. Our set generator predicts the correct language count on unseen pairs without assuming the number of languages, but underperforms oracle Top-k in exact set accuracy. Our analysis identifies the key obstacles to robust CS-LID: oracle cardinality, threshold instability, language bias in CS training data, and the synthetic-to-real gap.

CommentsAccepted at IEEE SLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑