arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28879cs.SDcs.AI

拓宽音频问答的不确定性估计:方法、格式与输入

Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs

Aaron Isidore Grace, Weiran Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究比较多种不确定性估计方法,发现首词度量高效且有效,且音频证据对不确定性影响远大于问题文本,为音频问答提供了可靠基线。

中文摘要 AI 辅助

音频语言模型可能产生缺乏音频支持的自信答案,这促使我们采用不确定性估计来识别不可靠的响应。我们比较了基于概率、基于采样、自我验证、证据和对比度量,涵盖四个开放权重模型和五个音频问答基准。在多项选择评估中,首词度量总体最强,top-1概率的平均AUROC为0.740,而十样本离散语义熵为0.708,且无需额外模型调用。在四个基准上,从多项选择转向开放式评估使平均准确率从57.6%降至36.6%,但不确定性仍可预测错误:语义熵、最大词元熵和语义一致性分别达到0.697、0.694和0.693的平均AUROC。为测试不确定性是否反映回答问题可用的证据,我们进行输入消融,移除音频或问题。在top-1置信度、熵和基于采样的度量中,移除音频使错误检测AUROC平均降低0.101,而移除问题仅降低0.010。综上,这些结果建立了高效的不确定性基线,并表明音频语言模型中的不确定性在很大程度上更依赖于可用的音频证据,而非问题文本。

英文摘要

Audio-language models can produce confident answers unsupported by the audio, motivating uncertainty estimates that identify unreliable responses. We compare probability-based, sampling-based, self-verification, evidential, and contrastive measures across four open-weight models and five audio QA benchmarks. In multiple-choice evaluation, first-token measures are strongest overall, with top-1 probability achieving a mean AUROC of .740, compared with .708 for ten-sample discrete semantic entropy, while requiring no additional model calls. Across four benchmarks, shifting from multiple-choice to open-ended evaluation lowers mean accuracy from 57.6% to 36.6%, yet uncertainty remains predictive of errors: semantic entropy, maximum token entropy, and semantic agreement achieve mean AUROCs of .697, .694, and .693, respectively. To test whether uncertainty reflects the evidence available to answer the question, we perform input ablations that remove either the audio or the question. Across top-1 confidence, entropy, and sampling-based measures, removing audio reduces error-detection AUROC by .101 on average, compared with .010 when removing the question. Together, these results establish efficient uncertainty baselines and show that uncertainty in audio-language models depends substantially more on available audio evidence than on question text.

发表机构

  • University of Waterloo(滑铁卢大学)
  • University of Iowa(爱荷华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑