EvoAudio:音频理解的递归自我改进
EvoAudio: Recursive Self-Improvement for Audio Understanding
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Tsinghua University(清华大学)
- Tencent Hunyuan(腾讯混元)
- Amphion Technology Co., Ltd.(安菲翁科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
EvoAudio提出递归自我改进系统,闭环进化模型、波形、问题与难度,利用模型性能指导数据生成,经13轮在MMSU等基准上提升多模型性能达6.3点。
AI中文摘要:
音频语言模型对所说内容的理解远胜于对其声音特性的理解。缩小这一差距需要的不仅仅是数据。详细的声学标注成本高昂,来自更强模型的标签会继承其错误和局限,而固定数据无法随着学习者的进步而适应。因此,我们提出EvoAudio,一个用于音频理解的递归自我改进系统。据我们所知,它是首个在单一闭环中同时进化模型、波形、问题和难度的系统。EvoAudio利用当前模型的性能来设定下一轮训练数据的焦点和难度。随后,一个音频工具库构建问题,其答案源自音频的制作方式,从而在无需新的人工标注的情况下提供可验证的监督。强化学习更新模型,验证则决定其是否进入下一轮进化。在13轮中,EvoAudio在MMSU、MMAU-Pro和MMAR上改进了五个具有不同音频编码器和语言骨干的模型。它为每个骨干取得了最高平均分,将整体性能提升了最多6.3个百分点。改进在连续轮次中逐步展开,每个更强的模型都从下一轮开始。
英文摘要:
Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio, a recursive self-improvement system for audio understanding. To our knowledge, it is the first to evolve the model, waveforms, questions, and difficulty in one closed loop. EvoAudio uses the current model's performance to set the focus and difficulty of the next training data. A library of audio tools then constructs questions whose answers follow from how the audio was made, providing verifiable supervision without new human annotation. Reinforcement learning updates the model, and validation decides whether it enters the next evolution round. Across 13 rounds, EvoAudio improves five models with different audio encoders and language backbones on MMSU, MMAU-Pro, and MMAR. It achieves the highest average for every backbone, raising overall performance by up to 6.3 points. The improvement unfolds over successive rounds, with each stronger model starting the next round.