arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BiMamba2掩蔽离散单元预测用于野外无监督语音挑战赛的多语言语音表示

BiMamba2 Masked Discrete-Unit Prediction for Multilingual Speech Representation for Unsupervised Speech in the Wild Challenge

Prakriti Subedi, Howard Prioleau, Saurav K Aryal

arXiv 2609.28758首次发表:更新:

发表机构

Howard University(霍华德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于BiMamba2的掩蔽离散单元预测方法,在无标注多语言语音上训练,实现说话人聚类性能超越基线,并分析了评估差异。

AI 中文摘要

我们描述了提交给Interspeech 2026野外无监督语音(UPS)挑战赛的系统,该系统采用双向Mamba-2(BiMamba2)编码器,遵循HuBERT风格范式,通过掩蔽离散单元预测进行训练。该47.88M参数模型在MLCommons无监督人民语音数据集的67种语言、250小时语音上进行训练,不使用任何标注数据。训练目标结合了掩蔽k均值伪标签预测、语言识别监督和VICReg正则化。在官方评估中,系统在说话人聚类上取得了0.735的调整兰德指数,超过了四个基线。语言识别宏F1(0.073)和字符错误率(0.870)仍低于监督基线。我们分析了本地与官方在指标尺度和检查点排名上的差异,强调了分布内诊断在预测Dynabench探针结果方面的局限性。

英文摘要

We describe our submission to the Unsupervised Speech in the Wild (UPS) Challenge at Interspeech 2026, a bidirectional Mamba-2 (BiMamba2) encoder trained with masked discrete-unit prediction following the HuBERT-style paradigm. The 47.88M-parameter model is trained on 250 hours of speech across 67 languages from the MLCommons Unsupervised People's Speech dataset, with no labeled data. The objective combines masked k-means pseudo-label prediction with language identification supervision and VICReg regularization. On official evaluation, the system achieves an Adjusted Rand Index of 0.735, exceeding four baselines on speaker clustering. Language identification macro-F1 (0.073) and character error rate (0.870) remain below supervised baselines. We analyze a local-official discrepancy in metric scale and checkpoint ranking, highlighting limitations of in-distribution diagnostics for predicting Dynabench probe outcomes.

CommentsAccepted to Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑