发表机构
Institute for Artificial Intelligence and Data Science; University at Buffalo(人工智能与数据科学研究院; 布法罗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出用衰减区域群时延作为取证线索,通过随机森林、CNN等模型验证其可区分AI生成脉冲声与真实脉冲声,为AI生成音频的溯源提供了新方法。
AI 中文摘要
我们研究是否可通过群时延分析区分AI生成脉冲声与真实脉冲声。核心发现为:AI生成脉冲声的起始区域群时延分布近乎相同,但在晚期衰减区域的群时延行为存在可测量差异——衰减区域KL散度达0.322,而起始区域散度接近零(0.022);跨频带GD变异性作为单一特征的AUC约为0.720,基于9个衰减区域特征的随机森林(RF)在样本不相交评估下AUC达0.884;将群时延图作为独立2D输入输入CNN分类器,准确率达90%-94%,表明群时延携带大量判别信息;在生成器留一测试下,CNN与Transformer分类器的AUC差异较大(0.457-0.918),群时延RF在评估方法中平均留一准确率最高(66.7%),且未出现远低于随机水平的崩溃,尽管其平均AUC(0.731)低于CNN平均(0.762)与AST(0.772);对27种STFT配置的参数敏感性分析证实,RF的AUC保持稳定(0.700-0.847,标准差=0.035)。这些结果表明,衰减区域群时延可作为物理可解释的取证线索,补充基于幅度的分类器,但仍需更广泛的验证。
英文摘要
We investigate whether AI-generated impulsive sounds can be distinguished from real ones through group delay analysis. Our central finding is that AI-generated impulsive sounds show near-identical onset-region group-delay distributions but exhibit measurably different group-delay behavior in the late decay region: decay-region KL divergence reaches $0.322$ compared to near-zero onset divergence ($0.022$). Cross-band GD variability achieves single-feature AUC~=~0.720, and a Random Forest (RF) over nine decay-region features reaches AUC~$=$~0.884 under sample-disjoint evaluation. A group delay map used as a standalone 2D input to CNN classifiers achieves 90--94\% accuracy, demonstrating that group delay carries substantial discriminative information. Under generator hold-out, CNN and transformer classifiers show highly variable AUC (0.457--0.918). The group delay RF achieves the highest average hold-out accuracy among the evaluated methods ($66.7\%$) and avoids extreme below-random collapse, although its average AUC (0.731) is lower than CNN avg (0.762) and AST (0.772). Parameter sensitivity analysis across 27 STFT configurations confirms that the RF AUC remains stable (0.700--0.847, std~=~0.035). These results suggest that decay-region group delay can serve as a physically interpretable forensic cue that complements magnitude-based classifiers, while broader validation remains necessary.
Comments6 pages, 4 figures, 7 tables. Submitted to IEEE International Workshop on Information Forensics and Security (WIFS) 2026