FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation
FireRedAudio:一种具有解耦连续表示的通用音频语言模型,用于理解与生成
机构 * Xiaohongshu(小红书)
AI总结 FireRedAudio是首个公开的统一音频-语言模型,采用解耦连续表示,支持音频理解、多语种ASR、各类TTS及语音编辑,性能优于相关模型。
Comments 20 pages, 3 figures