韵律转文本:从低通滤波语音中预测文本
Prosody-to-Text: Predicting text from low-pass filtered speech
浏览论文内容
中文总结 AI 辅助
该研究针对被忽视的韵律转文本任务,微调Whisper模型从低通滤波语音预测文本,取得WER36%等结果,揭示低频语音与词汇内容的强关联,或可用于指导LLM文本生成。
中文摘要 AI 辅助
尽管从文本预测韵律是该领域已确立的任务,但相反的方向——预测符合给定韵律模式的文本——仍在很大程度上被忽视。我们认为这很遗憾,因为该相反方向可能会带来一些非常有趣的用例。因此,在本文中,我们迈出了韵律转文本方向的第一步,研究能从韵律模式中恢复多少原始句子。为此,我们仅使用12个最低梅尔频带(截止频率约为450Hz的低通滤波器)对Whisper模型进行微调,获得了惊人准确的结果:词错误率(WER)为36%,其中10%的话语被完美恢复,40%的话语词错误率在25%或更低。我们还发现,给定正确前缀时,下一个标记的预测准确率达79%。我们的结果表明,低频语音特征与词汇内容之间的关系比之前认为的要强得多,我们相信,对该主题投入更多关注可能会为新应用打开大门,例如使用韵律来指导现代大型语言模型(LLM)的文本生成。
英文摘要
While predicting prosody from text is an established task in the field, the opposite direction, predicting text that fits a given prosodic pattern, remains largely overlooked. We find this unfortunate, because this opposite direction could lead to some very interesting use cases. Therefore, in this paper, we make the first steps in the prosody-to-text direction by inves- tigating how much of the original sentence can be recovered from its prosodic pattern. To this end, we fine-tune the Whis- per model using only the 12 lowest Mel bins (low-pass filter with approximately 450Hz cutoff), and obtain surprisingly accurate results (WER 36%), with 10% of utterances be- ing recovered perfectly, and 40% of utterances having Word Error Rate at or below 25%. We also find that, given the correct prefix, the next token was predicted correctly in 79% of cases. Our results suggest that the relationship between low-frequency speech features and lexical content is much stronger than previously thought, and we believe that direct- ing more attention to this topic might open the door to new applications, such as using prosody to guide text generation of modern LLMs
发表机构
- Masaryk University(马萨里克大学)
- Faculty of Informatics, Masaryk University(马萨里克大学信息学院)
- Natural Language Processing Centre(自然语言处理中心)
机构由 AI 辅助整理,请以论文原文为准。