发表机构
Zuoyebang Education Technology(作业帮教育科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DuplexDrama是首个涵盖场景、全双工行为、表现性语音和声音事件的合成对话数据集,通过4阶段流水线构建,包含超2000小时音频,并发布6400个双语对话子集以推进全双工对话模型研究。
AI 中文摘要
我们提出了DuplexDrama,这是首个同时涵盖四个维度的合成口语对话数据集:(i)完整的人物角色和场景设置;(ii)三种全双工行为(打断、反馈语、不完整语句);(iii)具有与角色一致的情感标签的表现性语音;(iv)基于脚本的声音事件。DuplexDrama通过一个4阶段流水线构建;对脚本和合成音频的质量验证确认了其质量。我们已生成了超过2,000小时的音频数据,包含跨越13种人物角色和5个年龄段共64种音色的音色池;所有对话轮次中有3.8%至少携带一种全双工行为。该数据已通过内部全双工模型训练得到验证。我们将发布一个精选的子集,包含6,400个双语对话(800小时,中文约500小时+英文约300小时),以推进全双工口语对话模型的研究。数据样本可在我们的演示页面获取,LLM评判的评估提示将与数据集一同发布。
英文摘要
We present DuplexDrama, the first synthesized spoken dialogue dataset that simultaneously covers four dimensions: (i) complete persona and scenario settings; (ii) three full-duplex behaviors (interruption, backchannel, incomplete); (iii) expressive speech with persona-aligned emotion labels; and (iv) script-aware sound events. DuplexDrama is built via a 4-stage pipeline; quality validation on both scripts and synthesized audio confirms its quality. We have produced more than 2,000 hours audio data with a 64-voice timbre pool spanning 13 personas and 5 age buckets; 3.8% of all turns carry at least one full-duplex behavior. This data has been validated through internal full-duplex model training. We will release a curated subset of 6,400 bilingual dialogues (800 h, Chinese ~500 h + English ~300 h) to advance full-duplex spoken dialogue model research. Data samples are available at our demo page and LLM-judge evaluation prompts will be released with the dataset.
Comments5 pages, 5 figures, 5 tables, 19 references. Demo: https://dunjie5465.github.io/duplexdrama-demo/