目标之前的种子:重新审视低资源加尔瓦利语自动语音识别(ASR)的评估
Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR
浏览论文内容
中文总结 AI 辅助
针对低资源加尔瓦利语ASR评估,构建首个可复现多种子基准,发现多种子评估可区分真实增益与种子噪声,且w2v-BERT 2.0搭配标准CTC表现最优。
中文摘要 AI 辅助
在低资源方言典型的语料库规模下,单次运行的对比可能产生无法复现的结果。我们以加尔瓦利语为例,该语言是喜马拉雅中部一种资源匮乏的印度-雅利安语,我们在官方VAANI划分上构建了首个可复现的多种子ASR基准,包含每个种子的输出及显著性检验。重新检验看似合理的增益后,我们发现这些增益并不稳定:在种子级测试下,无论是Focal CTC还是带matra加权的目标函数,都未优于标准CTC;matra目标函数甚至未能减少其针对的错误;印地语到加尔瓦利语的迁移学习也未比直接微调带来增益。真正有效的是常规方案:w2v-BERT 2.0搭配标准CTC在5个种子上达到47.0%的词错误率(WER),优于更大规模的MMS-1B及同类模型;预训练设计而非参数数量决定性能,速度增强带来小幅且基本稳定的增益。对官方划分的多种子评估可区分真实增益与种子噪声。
英文摘要
At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first reproducible multi-seed ASR benchmark on the official VAANI splits, with per-seed outputs and significance testing. Re-examining plausible gains, we find them fragile: neither Focal CTC nor a matra-weighted objective beats standard CTC under seed-level testing, the matra objective fails to cut even its targeted errors, and Hindi-to-Garhwali transfer gives no gain over direct fine-tuning. What holds up is mundane: w2v-BERT 2.0 with standard CTC reaches 47.0% WER over five seeds, beating the larger MMS-1B and comparable models; pretraining design, not parameter count, drives performance, and speed augmentation gives a small, largely consistent gain. Multi-seed evaluation on official splits separates real gains from seed noise.
发表机构
- Thapar Institute of Engineering and Technology(塔帕尔工程技术学院)
- Ulster University(阿尔斯特大学)
机构由 AI 辅助整理,请以论文原文为准。