这首曲目里有多少AI?量化混合音乐混音中AI生成音轨的占比
How Much AI Is in This Track? Quantifying the Proportion of AI-Generated Stems in Hybrid Music Mixtures
浏览论文内容
中文总结 AI 辅助
本文将AI音乐检测重新表述为连续AI能量比例的回归问题,提出基于CNN的模型实现了0.076的MAE和0.85的R²,为现实音乐制作中的AI音乐检测提供了有前景的初步方案。
中文摘要 AI 辅助
AI生成音乐在音轨(stem)层面的应用日益广泛,音乐制作人会将合成鼓、贝斯线或人声与人类演奏的乐器整合在一起。然而,当前的AI音乐检测系统是二分类的,将曲目视为完全AI生成或完全人类演奏。在本文中,我们将AI音乐检测重新表述为关于连续AI能量比例α(取值范围为[0,1])的回归问题。我们提出一种方法,利用多轨音乐数据集组装人类演奏和AI重构音轨(通过神经音频编解码器获得)的混合曲目,每种内容类型的比例已知。通过这种方法,我们首先表明,在完全AI生成或人类演奏的曲目上训练的基于CNN的模型,作为二分类检测器准确率超过99%,但面对混合内容时,其输出会随AI音轨的能量贡献上升,成为一个有噪声且校准不当的估计器。我们对不同音轨的影响分析表明,检测灵敏度取决于乐器并反映其频率内容:鼓和吉他带有强烈的编解码器伪影特征,而人声和贝斯则较难检测。基于这些见解,我们训练了一个类似的基于CNN的模型来回归α,在同一流程的保留混合曲目上实现了MAE=0.076,R²=0.85。这些结果表明,回归表述是迈向现实音乐制作工作流程中AI音乐检测的初步有希望的一步。
英文摘要
AI-generated music is increasingly used at the stem level, with producers integrating synthetic drums, basslines, or vocals alongside human-performed instruments. However, current AI music detection systems are binary, treating tracks as either fully AI or fully human. In this paper, we reformulate AI music detection as a regression problem on a continuous AI energy ratio, alpha in [0, 1]. We propose a methodology that leverages a multi-track music dataset to assemble mixtures of human-performed and AI-reconstructed stems (obtained using a neural audio codec) with known proportions of each content type. Using this approach, we first show that a CNN-based model trained on fully AI-generated or human-performed tracks, which achieves >99% accuracy as a binary detector, when faced with mixed content, yields an output that rises with the AI stems' energy contribution, acting as a noisy and miscalibrated estimator. Our analysis of the influence of different stems shows that detection sensitivity depends on the instrument and reflects its frequency content: drums and guitar carry strong codec-artifact signatures, while vocals and bass are less detectable. Based on these insights, we train a similar CNN-based model for regression of alpha, achieving MAE = 0.076 and R^2 = 0.85 on held-out mixtures from the same pipeline. These results suggest that the regression formulation is an initial promising step towards AI-music detection in realistic music production workflows.
发表机构
- Universitat Pompeu Fabra(庞培法布拉大学)
- BMAT Licensing S.L.(BMAT授权有限公司)
机构由 AI 辅助整理,请以论文原文为准。