arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估广播监测场景下的AI生成音乐检测

Assessing AI-generated music detection in real-world broadcast monitoring

David López-Ayala, Fernando García de la Cruz, Pablo Zinemanas, Emilio Molina, Martín Rocamora

arXiv 2608.07359首次发表:更新:

发表机构

Universitat Pompeu Fabra; BMAT Licensing S.L.(庞培法布拉大学; BMAT授权有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出真实电视录制数据集BAMM,对比不同训练方式的CNN在三类场景下的AI生成音乐检测性能,发现现有CNN检测方法在真实广播场景存在显著领域差距,检测性能不足。

AI 中文摘要

AI生成音乐在广播媒体中的普及引发了透明度与公平补偿方面的担忧,但在真实广播条件下进行可靠检测的问题仍未解决。现有研究报告显示,该领域的检测性能大幅下降,但其评估仅局限于合成广播数据。为填补这一空白,我们推出BAMM(Broadcast AI-Music Monitoring,广播AI音乐监测),这是一个包含AI生成音乐与人类创作音乐的40小时真实电视录制数据集。我们在三个难度逐步提升的场景中对比了仅在干净数据上训练和在广播数据上训练的CNN变体:干净前景音乐(CFM)、合成电视广播(STB)、真实电视广播(RTB)。两种模型在CFM场景下均达到近乎完美的性能,但在合成广播条件下性能大幅下降;与干净训练相比,面向广播的训练提升了鲁棒性,但性能仍有限。在使用BAMM进行评估的RTB场景中,两种模型性能进一步下降,且AI生成音乐与人类创作音乐的得分存在大量重叠。这些结果凸显了关键的领域差距,表明当前基于CNN检测器的训练方法仍不足以实现广播监测场景下可靠的AI生成音乐检测。

英文摘要

The proliferation of AI-generated music in broadcast media raises concerns about transparency and fair compensation, but reliable detection under real broadcast conditions remains unresolved. Existing studies report substantial performance degradation in this domain, yet their evaluations are limited to synthetic broadcast data. To address this gap, we introduce BAMM (Broadcast AI-Music Monitoring), a 40-hour dataset of real-world television recordings containing AI-generated and human-made music. We compare clean-trained and broadcast-trained CNN variants across three progressively more challenging scenarios: Clean Foreground Music (CFM), Synthetic TV Broadcast (STB), and Real TV Broadcast (RTB). Both models achieve near-perfect performance on CFM but degrade substantially under synthetic broadcast conditions. Broadcast-oriented training improves robustness compared with clean training, although performance remains limited. On RTB, evaluated using BAMM, both models degrade further and show substantial score overlap between AI-generated and human-made music. These results expose a critical domain gap and show that current training approaches on CNN-based detectors remain insufficient for reliable AI-generated music detection in broadcast monitoring.

CommentsAccepted for ISMIR 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑