发表机构
The University of Melbourne(墨尔本大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对语言文档编制的音频转录瓶颈,提出无需代码的开源ASR工作流Easper,通过评估转录优先级策略,发现优先词汇丰富叙述并增加声学语音重复可更快提升转录质量。
AI 中文摘要
音频转录是语言文档编制中的关键瓶颈。尽管Whisper等多语种自动语音识别(ASR)模型提供了解决方案,但田野语言学家往往缺乏利用这些模型的专业知识。我们提出Easper,这是一个开源、无需代码的工作流,使语言学家能够直接通过云资源基于ELAN注释迭代微调ASR模型。部署ASR还会引发冷启动问题:确定首先转录哪些录音以启动准确模型。使用Easper,我们对三种瓦努阿图语言(比斯拉马语、纳夫桑语、恩古纳语)评估转录优先级策略。我们按录音会话微调模型,比较优先考虑声学清洁度与语言丰富度时的字符错误率轨迹。我们证明,优先考虑词汇丰富的叙述并增加声学语音重复,即使在嘈杂环境中,也能更快提升转录质量。
英文摘要
Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack the expertise to utilise them. We present Easper, an open-source, no-code workflow enabling linguists to iteratively fine-tune ASR models via cloud resources directly from ELAN annotations. Deploying ASR also raises a cold start problem: deciding which recordings to transcribe first to bootstrap an accurate model. Using Easper, we evaluate transcription prioritisation strategies on three Vanuatu languages (Bislama, Nafsan, Nguna). We fine-tune models by recording session, comparing Character Error Rate trajectories when prioritising acoustic cleanliness versus linguistic richness. We demonstrate that prioritising lexically rich narratives and increasing acoustic-phonetic repetition, even in noisy environments, leads to faster improvements in transcription quality.
CommentsAccepted in Interspeech 2026