arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27817cs.SDcs.CLeess.AS

针对已知任务音频-LLM评估的生成式音频调用审计

Auditing generative audio calls for known-task audio-llm evaluation

  • National Research Council Canada(加拿大国家研究委员会)

机构由 AI 辅助整理,请以论文原文为准。

Mengzhe Geng

中文总结 AI 辅助

本研究将生成式音频调用评估设为受控决策问题,通过VocalSound数据集实验发现,带生成式调用的选择器准确率达0.925,优于无调用对照,明确了生成式调用在音频-LLM已知任务评估中的边际价值。

中文摘要 AI 辅助

语音与音频大语言模型(LLM)的评估常通过对比波形提示与自动语音识别(ASR)转写文本的表现完成。对于已知闭集任务,这种对比混淆了两个因素:声学证据的获取、以及调用生成式音频模型的需求。我们将该区分评估为受控调用决策问题。对于每个样本,策略在以下选项中选择:保留转写文本标签、使用对比语言-音频预训练(CLAP)、音频频谱图Transformer(AST)或WavLM的编码器证据、调用Qwen2-Audio、Qwen2.5-Omni或MOSS-Audio;决定性消融实验移除所有生成式调用操作,同时固定选择器与开发协议。在VocalSound数据集上,转写文本的准确率为0.296,因此需要波形信息。然而,有监督的CLAP与WavLM对照模型在无生成式音频调用时,准确率分别达到0.850与0.854。带有生成式调用操作的选择器使用12.5%的调用,达到0.925的准确率,相比之下,匹配的无调用选择器的准确率为0.921(配对差值为0.004;95%置信区间[-0.025,0.033])。一致性与堆叠特征可提升较弱选择器的表现,但无法超越最强的无调用对照模型。对于已知任务的端点声明,相关指标是在已使用转写文本与编码器证据后,生成式调用的边际价值。

英文摘要

Speech and audio LLMs are evaluated by comparing waveform predictions with predictions from an automatic speech recognition (ASR) transcript. For fixed closed-set tasks, this conflates acoustic evidence with the need to invoke a generative audio model. We estimate incremental call value with matched selectors sharing pre-call evidence. Each policy may retain the transcript label, use a local encoder, or invoke a generative model; matched control removes generative actions but preserves pre-call evidence and development selection. On VocalSound, transcript-only accuracy is 0.296, while supervised CLAP and WavLM controls reach 0.850 and 0.854 without calls. Full selector reaches 0.925 at 12.5% calls versus 0.921 for matched No-call selector (difference 0.004; 95% CI [-0.025, 0.033]). Thus, results do not show a call gain after transcript and encoder evidence are available. Relevant quantity is incremental accuracy from allowing calls, not the waveform-transcript gap.

↑