arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将AI生成音乐与编辑音频区分作为难负样本鲁棒性任务

Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task

Alexandru-Stefan Morosanu, Valerian Cecan, Stefan-Daniel Achirei, Laura Erhan

arXiv 2608.14916首次发表:更新:

AI 中文总结

本研究将区分AI生成音乐与编辑音频作为难负样本鲁棒性任务,构建YouTube数据集训练PaSST检测器,视频级平衡准确率0.811,片段级AI生成F1为0.836、编辑为0.720,揭示两者频谱线索的重叠性。

AI 中文摘要

AI生成音乐检测器通常针对原始歌曲进行评估,但现实世界中的上传内容常经过混音、重新编码、变调或其他编辑处理。这些编辑后的版本构成了难负样本类别:它们并非AI生成,但可能引入类似合成音频指纹的频谱伪影。我们将此问题作为AI生成音乐检测的难负样本鲁棒性场景进行研究,聚焦于源自相同锚定歌曲的AI生成及编辑变体。我们构建了基于YouTube的数据集,包含AI生成、编辑及原始变体,仅将原始曲目作为参考,并训练一个AI与编辑音频的二分类检测器。音频被处理为10秒片段,以原始波形形式输入预训练的PaSST频谱图Transformer。为减少数据泄露,所有划分均按锚定歌曲进行。在保留的测试集上,最终视频级系统达到0.811的平衡准确率;片段级上,AI生成片段的F1分数为0.836,而编辑片段的F1分数较低,为0.720。结果表明,AI生成音乐除普通编辑外仍保留可检测的类指纹频谱线索,但编辑类较低的F1分数显示这些线索仍可能与编辑音频的伪影重叠。Grad-CAM可视化用于检查高置信度预测是否依赖于局部时频区域。

英文摘要

AI-generated music detectors are commonly evaluated against original songs, but real-world uploads are often remixed, re-encoded, pitch-shifted, or otherwise edited. These edited versions form a difficult negative class: they are not generated by AI, yet they may introduce spectral artifacts that resemble synthetic audio fingerprints. We study this problem as a hard-negative robustness setting for AI-generated music detection, focusing on AI-generated and edited variants derived from the same anchor songs. We compile a YouTube-based dataset of AI, edited, and original variants, using the original tracks only as references, and train a binary AI versus edited detector. Audio is processed as 10-second clips and passed as raw waveforms to a pretrained PaSST spectrogram transformer. To reduce leakage, all splits are performed by anchor song. On the held-out test set, the final video-level system achieves 0.811 balanced accuracy. At clip level, AI-generated clips reach an F1-score of 0.836, while edited clips reach a lower F1-score of 0.720. The results suggest that AI-generated music retains detectable fingerprint-like spectral cues beyond ordinary editing, but the lower edited-class F1-score shows that these cues can still overlap with artifacts from edited audio. Grad-CAM visualizations are used to inspect whether high-confidence predictions rely on localized time-frequency regions.

CommentsAccepted at the RobustifAI 2026 Workshop @IJCAI-ECAI 2026, Bremen

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑