发表机构
Institute for Language and Speech Processing, Athena R.C.(雅典研究与技术中心语言与语音处理研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对希腊语低资源场景,首次系统研究 Whisper 模型规模、多任务组合及两阶段适应,构建对齐数据集,实现 WER 27.2%,建立首个希腊语 ALT 基准。
AI 中文摘要
自动歌词转录(ALT)由于旋律变化、节奏不规则和伴奏干扰,仍然比语音识别更具挑战性。这在像希腊语这样的低资源语言中尤为突出,此前不存在针对希腊语的 ALT 基准。我们首次对 Whisper 适应于希腊语 ALT 进行了受控研究,探讨了模型规模效应、通过多任务训练在转录-翻译比例中的任务组合,以及两阶段从语音到歌唱的适应。我们还基于希腊音频数据集(GAD),利用源分离和 CTC 强制对齐,整理了一个片段级对齐的歌唱数据集。结果表明,规模扩大持续提升性能,而多任务学习主要作为较小容量模型的有益正则化器。Whisper Large-v3 的两阶段适应实现了 27.2% 的词错误率(WER),相比零样本基线有显著改进,建立了首个希腊语 ALT 基准。
英文摘要
Automatic Lyric Transcription (ALT) remains substantially more challenging than speech recognition due to melodic variability, rhythmic irregularity, and accompaniment interference. This is heightened in low-resource languages like Greek, where no prior benchmark for ALT exists. We present the first controlled study of Whisper adaptation for Greek ALT, investigating model scaling effects, task composition via multitask training in transcribe-translate ratios, and two-stage speech-to-singing adaptation. We also curate a segment-level aligned singing dataset based on the Greek Audio Dataset (GAD) using source separation and CTC forced alignment. Results show that scaling consistently improves performance, while multitask learning acts as a beneficial regularizer primarily for smaller-capacity models. The 2-stage adaptation in Whisper Large-v3 achieves a Word Error Rate (WER) of 27.2%, a significant improvement over zero-shot baselines, establishing the first Greek ALT benchmark.
CommentsAccepted at Interspeech 2026