arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AudioICL-Bench:大型音频语言模型上下文学习基准

AudioICL-Bench: A Benchmark for Large Audio Language Model In-Context Learning

Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang, Ke-Han Lu, Hung-Yi Lee

arXiv 2609.11252首次发表:更新:

发表机构

National Taiwan University; Massachusetts Institute of Technology; NTU AI-CoRE(国立台湾大学; 麻省理工学院; 国立台湾大学人工智慧核心研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出AudioICL-Bench基准,通过重采样规则隔离先验知识,评估大型音频语言模型的上下文学习能力,发现时间感知与多规则组合是主要瓶颈。

AI 中文摘要

上下文学习(ICL)为音频领域提供了无需训练的适应方法,因为为每种新条件标注数据成本高昂。然而,现有的音频ICL研究主要衡量任务识别,即演示仅提示预训练能力,而非任务学习,即必须仅从演示中推断出真正的新输入-标签映射。我们引入了AudioICL-Bench,一个诊断性基准,其每个情节的规则被重新采样,使得无法从先验知识中恢复正确答案。其九个任务沿两个轴组织,将必须从演示中学习的内容与必须从信号中感知的内容分开,从而能够将失败归因于任一来源。在五个大型音频语言模型中,最强的模型在感知容易时能轻松将任意声音绑定到新标签,但在时间测量和组合多个诱导规则时崩溃,揭示了两个主要能力边界:时间感知和多规则组合。

英文摘要

In-context learning (ICL) promises training-free adaptation for audio, where labeling every new condition is costly. Yet existing audio ICL studies largely measure Task Recognition, where demonstrations merely cue pre-trained capabilities, rather than Task Learning, where a genuinely new input-label mapping must be inferred from demonstrations alone. We introduce AudioICL-Bench, a diagnostic benchmark whose per-episode rules are resampled so that no correct answer is recoverable from prior knowledge. Its nine tasks are organized along two axes that separate what must be learned from demonstrations from what must be perceived in the signal, enabling failures to be attributed to either source. Across five Large Audio Language Models, the strongest models readily bind arbitrary sounds to new labels when perception is easy, but collapse on temporal measurement and composing multiple induced rules, revealing two primary capability boundaries: temporal perception and multi-rule composition.

CommentsAccepted by SLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑