发表机构
Perle(珀尔)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建涵盖11种音频分类方法的分层基准,在2242个音频片段的闭集声源识别任务中评估性能,发现Gemini-3.1-Pro-Preview表现最优,还分析了Gemini模型的思维链特性并给出方法选择指导。
AI 中文摘要
我们对11种音频分类方法进行基准测试:5种任务感知闭集大语言模型(LLM)(4种Gemini模型加上开源权重的Kimi-Audio-7B-Instruct)、4种固定词汇标记器(YAMNet、PANNs、Whisper-AT和SSLAM)、1种零样本音频-文本模型(CLAP)以及1种音频基础大语言模型(BAT)。我们在一项闭集声源识别任务上对这些方法进行评估,该任务涉及2242个音频片段,涵盖23个细粒度类别和11个大类。由于这些方法在接收任务的方式和输出评分方式上存在根本差异,我们将它们分为四个评估层级而非单一排行榜,报告每个层级的宏查准率、宏召回率、宏F1值和漏检率。表现最佳的模型Gemini-3.1-Pro-Preview达到85.6%的大类级F1值和56.7%的细粒度F1值;Kimi-Audio因模型规模具有竞争力,达到67.5%的大类级F1值和32.9%的细粒度F1值,但对1.6%的样本未能给出答案;SSLAM和CLAP在不查看候选列表的情况下,大类级表现与最佳闭集模型相当或更优,但细粒度级表现落后。我们对Gemini模型的8968条响应进行思维链分析,发现响应长度无法预测准确率,所谓“整体判断优于详细分析”的效应更适合解释为难度混淆因素,且错误答案被自信陈述的比例达92%至100%。我们报告了所有11种方法的完整每类混淆矩阵和指标,确定了导致不同粒度间准确率损失的结构性错误模式,并为选择这些方法系列提供实用指导。
英文摘要
We benchmark eleven audio classification methods: five task-aware closed-set LLMs (four Gemini models plus open-weight Kimi-Audio-7B-Instruct), four fixed-vocabulary taggers (YAMNet, PANNs, Whisper-AT, and SSLAM), a zero-shot audio-text model (CLAP), and an audio-grounded LLM (BAT). We evaluate them on a closed-set sound-source identification task over 2,242 clips spanning 23 fine-grained classes and 11 categories. Since these methods differ fundamentally in how they receive the task and how outputs are scored, we group them into four evaluation tiers rather than one leaderboard, reporting macro Precision, Recall, F1, and false-negative rate per tier. The best model, Gemini-3.1-Pro-Preview, reaches 85.6 percent category-level F1 and 56.7 percent fine-grained F1. Kimi-Audio is competitive for its size, reaching 67.5 percent category-level F1 and 32.9 percent fine-grained F1, but fails to answer 1.6 percent of samples. SSLAM and CLAP match or exceed the best closed-set model at the category level without seeing the candidate list, but fall behind at the fine-grained level. Analyzing the Gemini models' chain-of-thought across 8,968 responses, we find that response length does not predict accuracy, an apparent "holistic judgment beats detailed analysis" effect is better explained as a difficulty confound, and wrong answers are stated confidently 92 to 100 percent of the time. We report full per-class confusion matrices and metrics for all eleven methods, identify the structural error modes behind most of the accuracy loss between granularities, and give practical guidance for choosing among these method families.