DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
DreamAudio: 基于扩散模型的定制化文本到音频生成
机构 * School of Computer Science and Electronic Engineering, University of Surrey(Surrey大学计算机科学与电子工程学院) ; Department of Informatics, King’s College London(伦敦国王学院信息学院) ; Seed Group, ByteDance Inc.(字节跳动Seed团队)
AI总结 本文提出DreamAudio,通过参考音频样本生成定制化音频,提升细粒度音频控制能力,实验显示其在定制生成任务中表现优异。
Comments Lastest arxiv version. Accepted by IEEE/ACM Transactions on Audio, Speech, and Language Processing. Demos are available at https://yyua8222.github.io/DreamAudio_demopage/