DuraS2ST:面向时长对齐的语音到语音翻译的思维链与强化学习
DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation
浏览论文内容
中文总结 AI 辅助
DuraS2ST提出时长对齐推理框架,利用思维链规划与强化学习优化,在CVSS-T上平衡翻译质量与时长一致性,优于现有基线。
中文摘要 AI 辅助
在视频配音等时间敏感应用中,语音到语音翻译(S2ST)不仅需要语义保真和说话人保留,还需要严格的时长一致性,以避免音视频错位。然而,现有的S2ST系统大多在没有显式时间规划的情况下生成目标语音,使得时长控制成为一个未解决的挑战。我们提出了DuraS2ST,一个时长对齐的推理框架,它使单个语音语言模型能够首先生成显式的思维链(CoT)来规划目标措辞和音素长度,然后合成相应的语音标记。为支持这一范式,我们构建了DuraSet-440K,一个高质量的时长对齐CoT语料库,用于监督初始化。我们进一步使用多模态多维强化学习优化模型,采用时长边际奖励(Duration Margin Reward)来平衡翻译质量和时长一致性,并使用模态感知奖励归因(Modality-Aware Reward Attribution)将奖励分配给适当的标记跨度。在CVSS-T上的实验表明,DuraS2ST在翻译质量和时长一致性之间取得了强劲的平衡,优于具有竞争力的开源和商业基线。项目页面:此https URL。
英文摘要
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- Microsoft Corporation(微软公司)
机构由 AI 辅助整理,请以论文原文为准。