arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知道何时不回答:音乐音频-语言模型的弃权伪集成方法

Knowing When Not to Answer: Pseudo-Ensembles for Abstention in Music Audio-Language Models

Aanya Maheshwari, Vatsal Raina

arXiv 2609.04362首次发表:更新:

发表机构

Jumeirah College; Apta AI; Spark AI Research(朱美拉学院; 阿普塔人工智能公司; 火花人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对音乐音频-语言模型无法判断自身未知答案的问题,提出基于预训练模型扰动输入的伪集成方法,提升了准确率与不确定性度量效果,且无需重新训练。

AI 中文摘要

音乐音频-语言模型的评估几乎完全基于多项选择题的准确率。这种评估协议迫使模型必须选定一个选项,因此侥幸猜中与真正的音乐理解无法区分。目前缺少的是一种方法,能够判断模型何时不知道答案,从而让模型可以弃权(不执行)而非猜测。常规解决方案是使用独立训练的模型集成,但在此场景下成本过高,因此仅能依赖单一预测分布的熵作为可用的置信度信号。我们改用预训练模型构建伪集成,通过以不改变正确答案的方式扰动其输入,再对各选项的结果分布取平均。我们的主要构建方式是简单打乱候选答案的呈现顺序;同时也研究了基于损坏音频和交换选项标签构建的集成。伪集成可为每个问题提供多个预测分布,因此支持基于集成的全部不确定性度量(期望分布的熵、期望熵,以及二者之差即互信息),而非仅依赖熵。在MuChoMusic数据集上评估TinyMU模型时,我们发现对4种选项顺序取平均可将准确率从55.7%提升至59.2%,且所得不确定性度量对模型错误的排序效果优于单遍熵基线,将错误保留曲线下面积从0.293降至0.261。所有操作仅需少量额外的前向传播,无需重新训练,这使得弃权对于紧凑的音乐音频-语言模型而言切实可行。

英文摘要

Music audio-language models are evaluated almost entirely by accuracy on multiple-choice questions. This protocol forces the model to commit to an option, so a lucky guess looks the same as real musical understanding. What is missing is a way to tell when the model does not know the answer, so that it can abstain instead of guessing. The usual solution, an ensemble of independently trained models, is far too expensive here, which leaves the entropy of a single predictive distribution as the only available confidence signal. We instead build pseudo-ensembles from one pretrained model by perturbing its input in ways that cannot change the correct answer, then averaging the resulting distributions over the options. Our main construction simply shuffles the order in which the candidate answers are presented; we also study ensembles built from corrupted audio and from swapped option labels. A pseudo-ensemble gives several predictive distributions per question, so it supports the full family of ensemble-based uncertainty measures (entropy of the expected distribution, expected entropy, and their difference, the mutual information) rather than entropy alone. Evaluating TinyMU on MuChoMusic, we find that averaging over four option orderings raises accuracy from 55.7% to 59.2%, and that the resulting uncertainty measures rank the model's errors better than the single-pass entropy baseline, reducing the area under the error retention curve from 0.293 to 0.261. All of this costs a few extra forward passes and no retraining, which makes abstention practical for compact music audio-language models.

Comments11 pages, 4 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑