arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FireRedAudio:一种具有解耦连续表示的通用音频语言模型,用于理解与生成

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

Feiyu Shen, Fenglong Xie, Junjie Li, Kun Xie, Lei Xie, Xu Tang, Xuelong Geng, Yan Jia, Yao Hu, Yichen Han, Yichen Wu, Ziqi Dai, Junjie Chen, Kai Huang, Manzhen Wei, Yixuan Li

arXiv 2608.24168首次发表:更新:

发表机构

Xiaohongshu(小红书)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FireRedAudio是首个公开的统一音频-语言模型,采用解耦连续表示,支持音频理解、多语种ASR、各类TTS及语音编辑,性能优于相关模型。

AI 中文摘要

统一音频模型需识别并理解语言、副语言及环境信息,同时支持语音合成与编辑。核心挑战在于表示:理解任务青睐适配长上下文建模的紧凑特征,而语音生成则需可重建、保留细粒度声学细节的特征。我们推出FireRedAudio,这是一种通用音频语言模型,采用共享的90亿参数大语言模型(LLM)。据我们所知,它是首个公开披露的统一音频-语言模型,在单个可训练自回归LLM内提供用于理解和生成的分离连续输入表示。待识别或分析的音频由专用音频编码器处理,而用于生成的语音输入则采用基于RedAE的通路。该LLM可直接生成文本,或对流匹配DiT进行条件设置以生成连续声学隐变量。通过渐进式多任务训练,FireRedAudio支持自动语音识别(ASR)与音频理解,后者可处理长达1小时的录音;还支持零样本文本转语音(TTS)、指令式TTS,以及语义与声学语音编辑。其对长音频的结构化组织达到秒级时间戳精度。经综合评估,FireRedAudio在音频理解与多语种ASR中取得具竞争力或领先的性能,在零样本TTS中展现出色的内容准确性与说话人保留能力,在指令式TTS中实现领先的指令遵循能力,且在语义与声学语音编辑上较Ming-UniAudio-Edit取得显著提升。这些结果表明,解耦连续输入表示可在中等规模模型中实现音频理解与连续隐变量语音生成的统一,相关代码可在指定URL获取。

英文摘要

A unified audio model must recognize and understand linguistic, paralinguistic, and environmental information while supporting speech synthesis and editing. A key challenge is representation: understanding favors compact features suited to long-context modeling, whereas speech generation requires reconstructible features that preserve fine-grained acoustic detail. We introduce FireRedAudio, a general-purpose audio language model with a shared 9B-parameter LLM. To the best of our knowledge, it is the first publicly disclosed unified audio-language model to provide separate continuous input representations for understanding and generation within a single trainable autoregressive LLM. Audio to be recognized or analyzed is processed by a dedicated Audio Encoder, while speech inputs for generation use a RedAE-based pathway. The LLM directly generates text or conditions a flow-matching DiT to produce continuous acoustic latents. Through progressive multitask training, FireRedAudio supports ASR and audio understanding, with the latter extending to recordings of up to one hour, as well as zero-shot TTS, Instruct TTS, and semantic and acoustic speech editing. Its structured organization of long-form audio achieves second-level timestamp accuracy. Across comprehensive evaluations, FireRedAudio achieves competitive or leading performance in audio understanding and multilingual ASR, strong content accuracy and speaker preservation in zero-shot TTS, leading instruction following in Instruct TTS, and substantial improvements over Ming-UniAudio-Edit in both semantic and acoustic speech editing. These results demonstrate the viability of decoupled continuous input representations for unifying audio understanding and continuous-latent speech generation in a model of moderate scale. Our code is available at https://github.com/FireRedTeam/FireRedAudio.

Comments20 pages, 3 figures. In this revision, the author list is ordered alphabetically by given name and an author-contribution statement is added; the technical content is unchanged

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑