AI 中文总结
该研究针对AI生成音乐检测器在简单音频操作下鲁棒性不足的问题,提出频率缩放不变的检测流水线,结合log-STFT重映射等技术,实现对速度修改等攻击的防御,还兼具可解释性。
AI 中文摘要
AI音乐生成器会留下由其架构决定的可预测频谱伪影。现有检测器利用这些伪影对原始生成曲目实现近乎完美的准确率,但在简单音频操作(如速度修改或音高偏移)下性能会崩溃。我们通过引入频率缩放不变的检测流水线解决这一开放性鲁棒性问题,旨在从设计层面防御此类攻击。我们的方法通过对数短时傅里叶变换(log-STFT)重映射将音频映射到对数频率轴,单个学习到的互相关滤波器结合最大池化在推理时提供平移不变性。训练使用混合损失,联合监督二元检测与伪影峰值定位,并对边界权重进行正则化。由于速度变化的鲁棒性是内置设计的,该检测器还具有可解释性:它同时输出二元决策和所应用速度变化因子的估计值。
英文摘要
AI music generators leave predictable spectral artifacts determined by their architecture. Existing detectors exploit these artifacts with near-perfect accuracy on raw generated tracks, but their performance collapses under simple audio manipulations, such as speed modification or pitch shifting. We address this open robustness problem by introducing a frequency-scaling-invariant detection pipeline that aims to prevent this kind of attack by design. Our method maps audio onto a log-frequency axis via a log-STFT remapping. A single learned cross-correlation filter, combined with max-pooling, provides shift invariance at inference time. Training uses a hybrid loss that jointly supervises binary detection and artifact-peak localization, regularizing boundary weights. Because robustness to speed change is built in by design, the detector is also interpretable: it outputs both a binary decision and an estimate of the applied speed-change factor.
CommentsProceedings of the 27th ISMIR Conference, Abu Dhabi, UAE, November 08-12, 2026