arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ArtifactBench:分布偏移下AI生成音乐检测器的谱系感知评估

ArtifactBench: Lineage-Aware Evaluation of AI-Generated Music Detectors under Distribution Shift

Heewon Oh

arXiv 2609.23550首次发表:更新:

发表机构

Intrect(Intrect)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出谱系感知评估套件ArtifactBench,在分布偏移下系统评估AI音乐检测器,揭示聚合分数掩盖的生成器与真实域性能差异。

AI 中文摘要

AI生成的音乐检测器通常使用基准测试上的聚合分数进行比较,而这些基准的训练重叠、生成器谱系、来源出处和音频变换历史仅部分可观测。本文介绍了ArtifactBench,一个谱系感知的评估套件,用于衡量检测器在生成器家族和版本、真实音乐域、收集队列偏移以及推理覆盖范围上的行为。该基准通过内容身份对源录音及其派生变体进行分组,将校准与最终测试分离,独立于分类错误记录推理失败,并在聚合指标之外报告带有不确定性的源级性能。我们在版本固定的通用协议下评估了多个公开可用的检测器,并考察了泄漏控制、队列可用性、阈值策略和模型特定缺失如何改变测量性能和模型排名。在562轨道的通用成功测试交集上,ArtifactNet获得了0.982的AUROC和0.918的平衡准确率,而公开的Deezer检测器为0.761/0.776;SpecTTTra和CLAM在此偏移队列下AUROC低于0.30。这些结果还揭示了聚合分数单独无法掩盖的实质性生成器和真实域偏移。

英文摘要

AI-generated music detectors are commonly compared using aggregate scores on benchmarks whose training overlap, generator lineage, source provenance, and audio-transformation history are only partially observable. This paper introduces ArtifactBench, a lineage-aware evaluation suite for measuring detector behavior across generator families and versions, real-music domains, collection-cohort shift, and inference coverage. The benchmark groups source recordings and their derived variants by content identity, separates calibration from final testing, records inference failures independently from classification errors, and reports source-level performance with uncertainty in addition to aggregate metrics. We evaluate multiple publicly available detectors under a version-pinned common protocol and examine how leakage control, cohort availability, threshold policy, and model-specific missingness alter measured performance and model ranking. On the 562-track common-success test intersection, ArtifactNet obtains 0.982 AUROC and 0.918 balanced accuracy, compared with 0.761/0.776 for the public Deezer detector; SpecTTTra and CLAM fall below 0.30 AUROC under this shifted cohort. These results also expose substantial generator- and real-domain shifts that aggregate scores alone conceal.

Comments9 pages, 1 figure, 6 tables. Related to ArtifactNet (arXiv:2604.16254); this work contributes the benchmark and evaluation protocol rather than a new detector

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑