揭示音频分类器中的捷径学习:通过发现时间解释中的重复概念
Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations
- Departamento de Computación, FCEyN, UBA(布宜诺斯艾利斯大学精确与自然科学学院计算系)
- ICC, CONICET-UBA(阿根廷国家科学与技术研究理事会-布宜诺斯艾利斯大学跨学科研究中心)
- Music and Audio Research Lab, New York University(纽约大学音乐与音频研究实验室)
- Integrated Design & Media, New York University(纽约大学综合设计与媒体系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出流水线,通过发现时间解释中的重复概念揭示音频分类器中的捷径学习,并验证其有效性。
AI中文摘要:
机器学习数据集中事件之间的相关性可能导致捷径学习,即模型基于相关事件的存在来预测目标事件,而非基于目标事件本身的特征。当这些相关性是虚假的——源于数据收集伪影——模型在实践中可能表现不佳。我们提出一个流水线,通过发现音频分类器时间解释中的重复概念来揭示捷径学习。具体而言,我们隔离解释分类器决策的音频片段,用大型音频语言模型集成对这些片段进行字幕描述,并使用大型语言模型提取重复概念。所得概念可由人类审计,以揭示潜在的捷径学习。我们使用从AudioSet Strong策划的数据集评估我们的框架,控制虚假相关性的存在与否。结果表明,该方法可靠地揭示了学习到的捷径,例如模型依赖“笑声”的存在来预测“掌声”。
英文摘要:
Correlations between events in machine learning datasets may result in shortcut learning, where models learn to predict the target event based on the presence of a correlated event. When these correlations are spurious -- arising from data collection artifacts -- models are likely to perform poorly in practice. We propose a pipeline to uncover shortcut learning in audio classifiers by discovering recurring concepts in their temporal explanations. Specifically, we isolate audio segments that explain classifier decisions, caption them with an ensemble of Large Audio-Language Models, and use a Large Language Model to extract recurring concepts. The resulting concepts can be audited by humans to uncover potential shortcut learning. We evaluate our framework using datasets curated from AudioSet Strong, controlling for the presence or absence of spurious correlations. Results show that this approach reliably uncovers learned shortcuts, such as the model relying on the presence of "laughter" to predict "applause".