发表机构
Howard University(霍华德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于BiMamba2的掩蔽离散单元预测方法,在无标注多语言语音上训练,实现说话人聚类性能超越基线,并分析了评估差异。
AI 中文摘要
我们描述了提交给Interspeech 2026野外无监督语音(UPS)挑战赛的系统,该系统采用双向Mamba-2(BiMamba2)编码器,遵循HuBERT风格范式,通过掩蔽离散单元预测进行训练。该47.88M参数模型在MLCommons无监督人民语音数据集的67种语言、250小时语音上进行训练,不使用任何标注数据。训练目标结合了掩蔽k均值伪标签预测、语言识别监督和VICReg正则化。在官方评估中,系统在说话人聚类上取得了0.735的调整兰德指数,超过了四个基线。语言识别宏F1(0.073)和字符错误率(0.870)仍低于监督基线。我们分析了本地与官方在指标尺度和检查点排名上的差异,强调了分布内诊断在预测Dynabench探针结果方面的局限性。
英文摘要
We describe our submission to the Unsupervised Speech in the Wild (UPS) Challenge at Interspeech 2026, a bidirectional Mamba-2 (BiMamba2) encoder trained with masked discrete-unit prediction following the HuBERT-style paradigm. The 47.88M-parameter model is trained on 250 hours of speech across 67 languages from the MLCommons Unsupervised People's Speech dataset, with no labeled data. The objective combines masked k-means pseudo-label prediction with language identification supervision and VICReg regularization. On official evaluation, the system achieves an Adjusted Rand Index of 0.735, exceeding four baselines on speaker clustering. Language identification macro-F1 (0.073) and character error rate (0.870) remain below supervised baselines. We analyze a local-official discrepancy in metric scale and checkpoint ranking, highlighting limitations of in-distribution diagnostics for predicting Dynabench probe outcomes.
CommentsAccepted to Interspeech 2026