准备发言:对齐大语言模型以生成适合文本转语音(TTS)的文本
Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation
浏览论文内容
中文总结 AI 辅助
本研究将生成适合TTS的文本视为LLM的偏好对齐问题,引入CORA和Recipe数据集,提出含多部分的评估套件,对比FaST框架与基线,发现FaST在TTS友好性与有用性间权衡最优,且指标间存在强相关。
中文摘要 AI 辅助
当前大语言模型(LLMs)主要针对书面文本进行优化,常生成语法正确且有用的输出,但不适合通过文本转语音(TTS)进行口语化传递。本研究探讨如何让LLMs原生生成适合TTS的文本,将其视为偏好对齐问题:不依赖下游改写模块,而是直接对齐LLMs以生成针对口语化传递优化的文本。我们引入两个覆盖不同目标领域的偏好数据集CORA和Recipe,包含成对的适合TTS与不适合TTS的响应。我们进一步提出一套评估套件,结合基于模式的启发式指标、TTS→ASR评估流程以及带有人类评判者的MUSHRA聆听研究。我们的实验对比了近期提出的特征感知采样与调优(FaST)框架——该框架利用可解释特征而非黑盒奖励模型——与一系列对齐基线在适合TTS的生成任务上的表现。值得注意的是,我们发现FaST在各种设置下实现了TTS友好性与有用性之间的最佳权衡。我们还确定了不同指标之间存在强相关性,凸显了通过高效启发式方法可靠评估TTS友好性的能力。
英文摘要
Current Large Language Models (LLMs) are primarily optimized for written text, often producing outputs that are grammatically correct and helpful yet poorly suited for spoken delivery via Text-to-Speech (TTS). In this work, we study how to make LLMs natively generate TTS-friendly text, which we frame as a preference alignment problem: instead of relying on downstream rewriting modules, we directly align LLMs to generate text optimized for spoken delivery. We introduce two preference datasets spanning different target domains, CORA and Recipe, which contain paired TTS-friendly and TTS-unfriendly responses. We further propose an evaluation suite combining a pattern-based heuristic metric, a TTS$\to$ASR evaluation pipeline, and a MUSHRA listening study with human judges. Our experiments compare the recently proposed Feature-aware Sampling and Tuning (FaST) framework -- leveraging interpretable features instead of a black-box reward model -- against an array of alignment baselines on the TTS-friendly generation task. Notably, we found that FaST achieves the best overall tradeoff between TTS-friendliness and helpfulness across various settings. We also identified a strong correlation between our different metrics, highlighting the ability to reliably assess TTS-friendliness via an efficient heuristic.
发表机构
- NAVER LABS Europe(NAVER欧洲实验室)
机构由 AI 辅助整理,请以论文原文为准。