发表机构
Graduate Institute of Communication Engineering, National Taiwan University, Taiwan; Artificial Intelligence Center of Research Excellence (AI-CoRE), National Taiwan University, Taiwan(国家台湾大学通讯工程研究所; 国家台湾大学人工智能卓越研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究音频-语言嵌入模型在处理否定时的问题,引入NegEval-Audio框架将数据集转换为否定感知任务,发现模型在否定情况下性能大幅下降,肯定偏差是缺陷,需明确否定感知训练目标。
AI 中文摘要
诸如CLAP之类的音频-语言嵌入模型在匹配当前声音事件方面得到广泛评估,但在否定方面很少。我们表明,仅肯定性评估隐藏了一个关键限制:这些模型无法编码否定的声音概念,将肯定和否定的字幕映射到几乎相同的表示。为揭示这一盲点,我们引入NegEval-Audio框架,将现有数据集转换为两个否定感知任务,检索否定和多项选择否定(MCQ-Neg),以探究模型能否区分出现和未出现的事件。在AudioCaps和Clotho数据集上,否定情况下性能急剧下降,否定类型的MCQ准确率远低于随机水平,即使对于基于多模态LLM的嵌入模型,这种失败仍然存在。虽然一种无需训练的引导方法改善了MCQ-Neg,但对检索否定的提升微乎其微。这表明肯定偏差是表示几何中的一个基本缺陷,需要明确的否定感知训练目标。
英文摘要
Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated sound concepts, mapping affirmative and negated captions to nearly identical representations. To expose this blind spot, we introduce NegEval-Audio, a framework that converts existing datasets into two negation-aware tasks, Retrieval-Neg and Multiple-Choice Negation (MCQ-Neg), to probe whether models distinguish present from absent events. On AudioCaps and Clotho, performance degrades sharply under negation, with negation-type MCQ accuracy falling far below chance, and the failure persists even for a recent multimodal LLM-based embedding model. While a training-free steering method improves MCQ-Neg, it yields marginal gains for Retrieval-Neg. This indicates that affirmation bias is a fundamental flaw in the representation geometry, necessitating explicit negation-aware training objectives.
CommentsManuscript in progress