DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
DIFFA-2:一种实用的扩散大型语言模型用于通用音频理解
机构 * College of Computer Science, Nankai University(南开大学计算机科学学院) ; Meituan LongCat Interaction Team(美团LongCat交互团队)
专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);instruction tuning(abstract);preference optimization(abstract)
AI总结 DIFFA-2是一种基于扩散模型的实用大型音频语言模型,通过升级语音编码器和双适配器提升音频理解性能,且在实际训练预算下表现优异。