arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HaikuS2S:一种用于以诗歌形式回应的级联系统

HaikuS2S: A Cascaded System For Responding In Verse

Devangi Sharma, Sophia Judicke, Glenda Tan, Conrad Schaumburg, Shinji Watanabe

arXiv 2609.23951首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有语音合成难以生成结构化诗歌的问题,提出结合ASR、LLM和TTS微调的级联系统HaikuS2S,以改善俳句的韵律与声调对齐,同时保持情感相似性。

AI 中文摘要

表现性语音合成已通过韵律建模取得进展,然而生成结构化的诗歌语音(如俳句)仍然具有挑战性。先前关于韵律迁移的工作提升了表现力,而微调后的诗歌TTS(文本到语音)系统能够捕捉诗句的语调。然而,这些模型并未建模俳句的5-7-5音节结构或行尾停顿。我们提出了一种级联系统HaikuS2S,结合了ASR(自动语音识别)、LLM(大语言模型)生成的俳句,以及在散文和自定义俳句数据集上进行TTS微调。我们的评估聚焦于情感相似性、语音质量和韵律对齐。在实验中,我们发现我们的韵律和声调对齐在微调系统中显著改善,尤其是同时基于一般诗歌和俳句训练的系统。我们还观察到所有系统在情感相似性得分上保持相近。

英文摘要

Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior work on prosody transfer improves expressiveness, and fine-tuned poetry TTS (text-to-speech) systems capture verse intonation. However, these models do not model haiku's 5-7-5 syllable structure or line-ending pauses. We present a cascaded system, HaikuS2S, combining ASR (automatic speech recognition), LLM (large language model)-generated haiku, and TTS fine-tuning on both prose and custom haiku datasets. Our evaluation focuses on emotion similarity, speech quality, and prosody alignment. In our experiments, we see that our prosody and tonal alignment improve significantly with our fine-tuned systems, particularly the one trained on both general poetry and haiku. We also see that we maintain similar emotion similarity scores across all systems.

CommentsAccepted to SLT 2026, Demo Track. 5 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑