重新思考语音基础模型微调:更好的监督微调还是更好的匹配?
Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?
浏览论文内容
中文总结 AI 辅助
研究语音基础模型微调,通过对3个SUPERB分类任务、9个预训练检查点的8种SFT变体进行系统研究,发现SFT结果强烈依赖预训练实例,顶级SFT方法常因检查点而异,下游增益多为实例和种子依赖的匹配,非普遍性能提升。
中文摘要 AI 辅助
监督微调(SFT)被广泛用于使自监督语音表示适应下游分类任务。在单个预训练检查点下观察到的微小增益通常被解释为方法层面的改进,即更高的可达到性能上限。我们表明此类结论并不总是可靠的,因为SFT结果强烈依赖于特定的预训练实例。我们对3个SUPERB分类任务进行了系统研究,评估了来自wav2vec~2.0、HuBERT和WavLM的9个预训练检查点的8种SFT变体,并在代表性基础规模模型上进行了多种子重复实验。我们发现统计上难以区分的顶级SFT方法的身份通常依赖于检查点,在预训练实例之间的可转移性有限。这些发现表明,许多报告的下游增益反映的是实例和种子依赖的诱导匹配,而不是普遍提高可达到的性能上限。
英文摘要
Supervised fine-tuning (SFT) is widely used to adapt self-supervised speech representations to downstream classification tasks. Small gains observed under a single pretrained checkpoint are often interpreted as method-level improvements, i.e., a higher attainable performance ceiling. We show that such conclusions are not always reliable because SFT outcomes depend strongly on the specific pretrained instance. We conduct a systematic study on 3 SUPERB classification tasks, evaluating 8 SFT variants across 9 pretrained checkpoints from wav2vec~2.0, HuBERT, and WavLM, with multi-seed repetitions on representative base-scale models. We find that the identity of the statistically indistinguishable top-group SFT recipe is often checkpoint-dependent, with limited transferability across pretrained instances. These findings suggest that many reported downstream gains reflect instance and seed dependent elicitation match, rather than universally improving the attainable performance ceiling.
发表机构
- Graduate School of Informatics, Kyoto University(京都大学信息学研究生院)
机构由 AI 辅助整理,请以论文原文为准。