arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缺失之音:音频-语言嵌入模型在处理否定时存在困难

The Sound of Absence: Audio-Language Embedding Models Struggle with Negation

Chun-Yi Kuan, Hung-yi Lee

arXiv 2607.12290首次发表:更新:

发表机构

Graduate Institute of Communication Engineering, National Taiwan University, Taiwan; Artificial Intelligence Center of Research Excellence (AI-CoRE), National Taiwan University, Taiwan(国家台湾大学通讯工程研究所; 国家台湾大学人工智能卓越研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究音频-语言嵌入模型在处理否定时的问题,引入NegEval-Audio框架将数据集转换为否定感知任务,发现模型在否定情况下性能大幅下降,肯定偏差是缺陷,需明确否定感知训练目标。

AI 中文摘要

诸如CLAP之类的音频-语言嵌入模型在匹配当前声音事件方面得到广泛评估,但在否定方面很少。我们表明,仅肯定性评估隐藏了一个关键限制:这些模型无法编码否定的声音概念,将肯定和否定的字幕映射到几乎相同的表示。为揭示这一盲点,我们引入NegEval-Audio框架,将现有数据集转换为两个否定感知任务,检索否定和多项选择否定(MCQ-Neg),以探究模型能否区分出现和未出现的事件。在AudioCaps和Clotho数据集上,否定情况下性能急剧下降,否定类型的MCQ准确率远低于随机水平,即使对于基于多模态LLM的嵌入模型,这种失败仍然存在。虽然一种无需训练的引导方法改善了MCQ-Neg,但对检索否定的提升微乎其微。这表明肯定偏差是表示几何中的一个基本缺陷,需要明确的否定感知训练目标。

英文摘要

Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated sound concepts, mapping affirmative and negated captions to nearly identical representations. To expose this blind spot, we introduce NegEval-Audio, a framework that converts existing datasets into two negation-aware tasks, Retrieval-Neg and Multiple-Choice Negation (MCQ-Neg), to probe whether models distinguish present from absent events. On AudioCaps and Clotho, performance degrades sharply under negation, with negation-type MCQ accuracy falling far below chance, and the failure persists even for a recent multimodal LLM-based embedding model. While a training-free steering method improves MCQ-Neg, it yields marginal gains for Retrieval-Neg. This indicates that affirmation bias is a fundamental flaw in the representation geometry, necessitating explicit negation-aware training objectives.

CommentsManuscript in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑