arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36921cs.SD

当能力无法组合:诊断大型音频-语言模型中的组合性差距

When Capabilities Fail to Compose: Diagnosing the Compositionality Gap in Large Audio-Language Models

发表机构国立台湾大学 · 台湾大学人工智能卓越研究中心 · 华硕开放云端基础设施软件中心
查看机构详情
  • National Taiwan University(国立台湾大学)
  • NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)(台湾大学人工智能卓越研究中心)
  • ASUS Open Cloud Infrastructure Software Center(华硕开放云端基础设施软件中心)

机构由 AI 辅助整理,请以论文原文为准。

Chien-Feng Liu, Chih-Kai Yang, Bo-Han Feng, Yu-Hsuan Li Liang, Hung-yi Lee, Cheng-Fu Chou

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过受控诊断实验,发现大型音频-语言模型在组合音频属性识别与下游任务时存在系统性能力组合差距,表现为准确率显著下降。

中文摘要 AI 辅助

大型音频-语言模型(LALMs)在单个音频任务上表现强劲,但这些能力能否被可靠地组合仍未被充分探索。我们对LALMs中的能力组合进行了一项受控的诊断性研究,要求模型整合音频属性识别、线索条件性片段选择以及下游的自动语音识别(ASR)或问答。我们构建了具有不同声学线索的双话语输入,以评估环境声音、性别和情感线索上的组合,并将ASR、数学问答和事实性问答作为下游任务。在四个开源LALMs中,组合性问答准确率在40个模型-任务-线索设置中的39个中下降,平均下降26.7个百分点。ASR表现出类似一致的退化,词错误率(WER)在40个设置中的39个中增加,平均增加28.5个百分点,而退化幅度因模型、线索类型和线索显著性而异。我们进一步通过输出格式、位置偏好和思维链(CoT)分析来探究这些失败。我们的研究揭示了拥有单个音频能力与可靠组合它们之间的系统性差距。

英文摘要

Large audio-language models (LALMs) perform strongly on individual audio tasks, but whether these capabilities can be reliably composed remains underexplored. We conduct a controlled diagnostic study of capability composition in LALMs, requiring models to integrate audio-attribute recognition, cue-conditioned segment selection, and downstream ASR or question answering. We construct two-utterance inputs with distinct acoustic cues to evaluate composition across environmental sound, gender, and emotion cues, with ASR, Math QA, and Factual QA as downstream tasks. Across four open-source LALMs, compositional QA accuracy decreases in 39 of 40 model-task-cue settings, by an average of 26.7 percentage points. ASR exhibits a similarly consistent degradation, with WER increasing in 39 of 40 settings by an average of 28.5 percentage points, while the magnitude of degradation varies across models, cue types, and cue salience. We further probe these failures through output format, positional preference, and chain-of-thought (CoT) analyses. Our study reveals a systematic gap between possessing individual audio capabilities and reliably composing them.

补充信息

↑