发表机构
University of Georgia(佐治亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到音乐模型数据归因缺乏可靠基准的问题,提出受控数据集TrueMuse,通过微调三个扩散模型构建,涵盖四种归因设置,系统评估现有黑盒方法,发现归因仍具挑战性,需更可靠通用的方法。
AI 中文摘要
文本到音乐生成模型在庞大的音乐数据集上进行训练,这产生了对数据归因方法的日益增长的需求,这些方法能够量化单个训练样本的贡献。然而,由于缺乏可靠的基准真值,现有的归因方法难以进行严格评估,这使得可靠地评估其实际有效性变得具有挑战性。为了解决这一空白,我们引入了TrueMuse,一个用于文本到音乐数据归因的受控数据集和基准。TrueMuse通过在三组精心策划的归因样本上微调三个基于扩散的文本到音乐模型而构建,这些样本已知包含在微调中,为评估提供了受控的归因目标。该基准涵盖四种归因设置,涉及旋律结构、音色特征、艺术家级风格特征和流派级共享模式,并包括133个属性、648个微调模型和95,456个生成样本,涵盖两种提示类型。利用TrueMuse,我们沿着四个维度系统评估了现有的黑盒归因方法:微调改进、提示类型难度、多任务训练和微调数据规模。我们的结果表明,归因仍然具有挑战性,现有方法在不同评估设置中表现出显著差异,这凸显了对更可靠和更通用的文本到音乐生成归因方法的需求。代码和数据集将在论文被接受后发布。
英文摘要
Text-to-music generation models are trained on massive music collections, creating a growing need for data attribution methods that can quantify the contribution of individual training samples. However, existing attribution methods are difficult to rigorously evaluate due to the lack of reliable ground truth, making it challenging to reliably assess their actual effectiveness. To address this gap, we introduce TrueMuse, a controlled dataset and benchmark for text-to-music data attribution. TrueMuse is constructed by fine-tuning three diffusion-based text-to-music models on carefully curated attribution samples, whose known inclusion in fine-tuning provides controlled attribution targets for evaluation. The benchmark covers four attribution settings, spanning melodic structure, timbral characteristics, artist-level stylistic signatures, and genre-level shared patterns, and includes 133 attributes, 648 fine-tuned models, and 95,456 generated samples across two prompt types. Using TrueMuse, we systematically evaluate existing black-box attribution methods along four dimensions: fine-tuning improvement, prompt-type difficulty, multi-task training, and fine-tuning data size. Our results show that attribution remains challenging, with existing methods exhibiting substantial variation across evaluation settings, highlighting the need for more reliable and generalizable attribution methods for text-to-music generation. Code and Dataset will be released upon acceptance.