arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

dots.tts.edit:基于连续自回归模型的精确可控语音编辑

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

Hankun Wang, Bohan Li, Shi Lian, Xiaoyu Gu, Jing Peng, Da Zheng, Yiwei Guo, Colin Zhang, Shuai Wang, Kai Yu

arXiv 2608.02673首次发表:更新:

发表机构

dots.tts Team(dots.tts团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出dots.tts.edit语音编辑工具,采用XML风格结构化指令接口,基于连续自回归TTS模型,在doteBench评估中表现优异,音频质量与现有系统相当。

AI 中文摘要

用于内容创作的语音编辑需要对编辑操作内容及应用位置进行精确控制。自由形式的自然语言为表达编辑请求提供了灵活的接口,但其模糊性可能导致预期操作、参数或目标区域未被明确说明。本研究提出了一种精确且明确的语音编辑接口:基于转录文本的结构化编辑指令,采用XML风格标签明确指定操作类型,并将其定位到转录文本的跨度或边界。该语义时间线避免了显式时间戳对齐,为组合编辑提供了可外部检查的约定。我们将该接口实例化为dots.tts.edit,这是一款基于连续自回归TTS基础模型改造的编辑器。四个代表性语音创作控制项通过文本、情感、韵律和停顿编辑,覆盖词汇内容、情感表达、音高与语速传递以及时间 phrasing(此处保留专业术语)。特定任务的数据管道构建了操作和范围受控的对,同时保留每个目标区域之外的源上下文。我们进一步推出doteBench,这是一个双语评估套件,用于测量四个控制项及其组合的精确指令遵循、局部保留和音频质量。实验表明,该模型在其五个编辑类别中实现了领先的整体指令遵循和局部保留,同时音频质量与现有开源系统相当。在三个Seed-TTS-Eval分片上,该模型在零样本TTS识别错误率和说话人相似度方面与基础模型相比差异可忽略不计。代码和模型将很快发布。

英文摘要

Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the intended operation, parameters, or target region underspecified. We study a precise and explicit interface for speech editing: a transcript-grounded structural edit instruction with XML-style tags explicitly specifies typed operations and localizes them to transcript spans or boundaries. This semantic timeline avoids explicit timestamp alignment and provides an externally inspectable contract for compositional edits. We instantiate the interface in dots$.$tts$.$edit, an editor adapted from the continuous autoregressive dots$.$tts foundation model. Four representative speech-creation controls cover lexical content, affective expression, pitch and speaking-rate delivery, and temporal phrasing through text, emotion, prosody, and pause editing. Task-specific data pipelines construct operation- and scope-controlled pairs while retaining source-derived context outside each target region. We further introduce doteBench, a bilingual evaluation suite that measures precise instruction following, local preservation, and audio quality across the four controls and their composition. Experiments show leading overall instruction following and local preservation across its five editing categories, while audio quality remains comparable to existing open-source systems. Across three Seed-TTS-Eval shards, the model shows negligible differences from the base model in zero-shot TTS recognition error rate and speaker similarity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑