发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出首个歌唱领域专家声乐教练反馈评估基准VocalCoachBench,通过实验发现现有音频-语言模型在细粒度问题标签识别等方面存在明显差距,推动音频-语言评估向分析性反馈发展。
AI 中文摘要
近期的音频-语言模型越来越多地在识别、描述和推理音频方面接受评估,但面向专家的应用需要不同的能力:生成基于输入的、能识别问题并提出纠正措施的反馈。我们推出VocalCoachBench,这是一个用于评估音频-语言模型在歌唱领域专家声乐教练反馈上表现的基准。VocalCoachBench包含18位专业声乐教练标注的515条录音,产生1056份专家提交内容和12051条原子级教练主张。它包含用于受控比较的同歌曲子集,以及用于在不同歌曲和录音条件下进行分段基础反馈的多样歌曲子集。为适应专家反馈的开放性,VocalCoachBench将确定性结构化目标与基于主张的自由形式诊断和纠正指导评估分开。人类标注分析显示,专家一致性随标签粒度变化显著,这促使了分层结构化指标和基于主张的开放性反馈评估的产生。对12种近期音频-语言模型的实验揭示了一个持续存在的差距:尽管模型能在自由形式反馈中比较表现并识别宽泛的问题领域,但前3项细粒度问题标签识别仍低于标签先验基线,严格诊断对齐率仍低于7%。据我们所知,VocalCoachBench提供了首个用于评估基于音频的歌唱专家反馈的公开测试平台,将音频-语言评估从描述推进到分析性反馈。
英文摘要
Recent audio-language models are increasingly evaluated on recognizing, describing, and reasoning about audio, but expert-facing applications require a different capability: producing feedback that identifies problems and suggests corrective actions grounded in the input. We introduce VocalCoachBench, a benchmark for evaluating audio-language models on expert vocal coaching feedback for singing. VocalCoachBench contains 515 recordings annotated by 18 professional vocal trainers, yielding 1,056 expert submissions and 12,051 atomic coaching claims. It comprises a same-song subset for controlled comparison and a diverse-song subset for segment-grounded feedback across varied songs and recording conditions. To accommodate the open-ended nature of expert feedback, VocalCoachBench sep- arates deterministic structured targets from claim-based assessment of free-form diagnosis and corrective guidance. Human annotation analysis shows that expert agreement varies strongly with label granularity, motivating hierarchical structured metrics and claim-based evaluation of open-ended feedback. Experiments with 12 recent audio-language models reveal a consistent gap: while models can compare performances and identify broad issue domains in free-form feedback, Top-3 fine-grained issue-label identification remains below label-prior baselines and strict diagnosis alignment stays below 7%. To our knowledge, VocalCoachBench pro- vides the first public testbed for evaluating audio-grounded expert feedback for singing, moving audio-language evaluation beyond description toward analytic feedback.