发表机构
NYU Courant Institute(纽约大学库朗研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出利用Whisper工具改进多语言视频语音转录,通过微调将七种语言的平均错误率从30%降至20%,以支持跨文化自动化工具的训练。
AI 中文摘要
在当今高度互联的跨国世界中,跨文化理解已变得日益重要。基于大语言模型(LLM)技术的成功正推动着自动化工具的开发,以帮助非母语人士在跨文化环境中取得成功。构建此类自动化工具通常依赖于利用自然环境中(in-the-wild)的文本、音频和视频数据。本文提出了改进从视频中生成多语言语音识别转录的技术,以更好地训练这些自动化工具。研究重点在于那些易于被跨文化工具构建者使用的流程和语音工具,无需深厚的语音处理专业知识。使用来自YouTube的公开视频和基于Whisper的工具,在七种语言(西班牙语、日语、韩语、普通话、土耳其语、俄语和希伯来语)中观察到平均转录错误率为30%。通过适量的微调数据,平均错误率可降低至20%,使得此类输出更适合下游处理。与这些视频相关的语音和元数据也已发布,供社区进一步优化这些实验。
英文摘要
Cross-cultural understanding has become increasingly important in today's highly connected, cross-national world. The success of LLM-based technologies is now driving the development of automated tools to aid understanding for nonnative people trying to succeed in cross-cultural environments. Building such automated tools is often done by leveraging in-thewild text, audio, and video data. This paper presents techniques for improving speech recognition-based transcript creation in multiple languages from videos to better train these automated tools. The focus is on processes and speech tools that can easily be used by cross-cultural tool builders without requiring deep speech processing expertise. Using publicly available videos from YouTube and Whisper-based tools, average transcription error rate across seven languages (Spanish, Japanese, Korean, Mandarin, Turkish, Russian, and Hebrew) of 30% are observed. With a modest amount of fine-tuning data, the average error rate can be reduced to 20% making such output much more usable for downstream processing. Speech and metadata associated with these videos that can be used by the community to further refine these experiments are released as well.
Comments7 pages, 2 figures, 5 tables