AI 中文总结
研究针对大型音频语言模型细粒度音频推理难题,提出无标签自我进化框架Audio-Zero,通过构建听觉自博弈游戏,利用无标签音频对比对让模型生成线索并推理,实验证明其能提升细粒度音频推理且保持广泛音频理解。
AI 中文摘要
大型音频语言模型(LALMs)在声学理解方面取得了快速进展,但在细粒度音频推理(如识别事件顺序、重复和持续时间)方面仍存在困难。现有训练后方法严重依赖昂贵的外部标签或仅提供粗略的语义信号。为弥补这一差距,我们引入了Audio-Zero,这是LALMs领域首个无标签自我进化框架,可改善细粒度听觉感知和推理。它从无标签音频对比对构建听觉自博弈游戏,模型生成线索并通过推理线索间不一致识别特殊听众,该游戏能提供可验证奖励。在TREA、MMAU Test-mini和MMAR上对Qwen2-Audio-7B-Instruct和Qwen2.5-Omni-7B的实验表明,Audio-Zero在保持广泛音频理解的同时改善了细粒度音频推理。进化和诊断分析进一步揭示,越来越细粒度的听觉描述自然地从游戏压力中出现。
英文摘要
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-Zero constructs an auditory self-play game from unlabeled audio contrast pairs: most players hear a reference audio, while one odd listener hears a subtle variant. The model first generates clues describing what it hears and then identifies the odd listener by reasoning over inconsistencies among clues. Since the odd listener is known by construction, the game provides verifiable rewards without any annotated answers. Experiments with Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA, MMAU Test-mini and MMAR show that Audio-Zero improves fine-grained audio reasoning while preserving broad audio understanding. Evolutionary and diagnostic analyses further reveal that increasingly fine-grained auditory descriptions emerge naturally from game pressure.