AI 中文总结
DCASE 2026任务5聚焦音频相关问答,用音频依赖过滤管道构建评估集。比赛分两轨道,众多团队参与。成均馆大学团队取得较好成绩,还分析了常见构建模块、测试方法及各系统普遍遗漏的项目。
AI 中文摘要
DCASE 2026任务5引入了音频相关问答(ADQA),旨在测试大型音频语言模型是否依据音频而非文本先验进行回答。通过音频依赖过滤(ADF)管道,结合静音音频探测、每个选项的困惑度、大语言模型常识检查和人工审核,去除仅可从文本解决的项目。通过的3000个项目构成ADQA-Bench评估集,涵盖音乐、语音和环境音频。首届比赛吸引了14个团队和36份提交作品,分两个参数计数轨道。成均馆大学的MOSS-Audio-8B-Thinking和Qwen3-Omni-30B组合以58.33%的总体准确率位居榜首,同一团队的仅MOSS配置在低于10B参数轨道以57.30%领先。在具有可比开发分数的30份提交作品中,隐藏评估集上的评估准确率平均下降11.91个百分点。最常见的构建模块包括MOSS-Audio-8B-Thinking主干、在AudioMCQ-StrongAC上的低秩适应(LoRA)微调以及偏好或强化学习目标。测试时,提示工程几乎普遍存在,多数或选择排列投票很常见。每个系统都遗漏了相同的233个评估项目。
英文摘要
DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Audio-Dependency Filtering (ADF) pipeline combines silent-audio probing, per-option perplexity, a large language model (LLM) commonsense check, and human review to remove items solvable from text alone. The 3000 items that pass form the ADQA-Bench evaluation set, spanning music, speech, and environmental audio. The inaugural edition draws 14 teams and 36 submissions across two tracks defined by total parameter count (up to 100B and under 10B). A Chung-Ang University ensemble of MOSS-Audio-8B-Thinking and Qwen3-Omni-30B reaches the top overall accuracy at \pct{58.33}, and a MOSS-only configuration from the same team leads the sub-10B track at \pct{57.30}. Across the 30 submissions with a comparable development score, evaluation accuracy falls by 11.91 percentage points (pp) on average (median 10.91\,pp) on the hidden evaluation split, which is designed to be harder than the development split. The most common building blocks are: the MOSS-Audio-8B-Thinking backbone (13 of 36 submissions), Low-Rank Adaptation (LoRA) fine-tuning on AudioMCQ-StrongAC, and preference or reinforcement-learning objectives -- Group Relative Policy Optimization (GRPO) in five teams, Group reward-Decoupled Normalization Policy Optimization (GDPO) in two. At test time, prompt engineering is near-universal, and majority or choice-permutation voting is common. Every system misses the same set of 233 evaluation items.