arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LILAC:一种幂等神经语音编解码器

LILAC: An Idempotent Neural Speech Codec

June Young Yi, Dongwook Lee, Jiheum Yeom, Sungroh Yoon

arXiv 2608.05727首次发表:更新:

AI 中文总结

本文提出具备结构幂等性的神经语音编解码器 LILAC,其码率0.75 kbit/s,在 LibriSpeech 和 LibriTTS-R 测试集上的 UTMOS 分别达4.14和4.24,性能与亚1 kbit/s 神经音频编解码器最优水平相当,解决了现有编解码器幂等性不足的问题。

AI 中文摘要

神经音频编解码器被广泛应用于语音生成与编辑领域,但现有神经音频编解码器不具备幂等性:在本文测试的12个基线系统中,每种配置在单次解码-重编码过程中平均至少会重写15%的 token,这为将神经音频编解码器用作管道中的 token 接口带来了问题,因为管道中可能会出现对解码后输出进行重编码的情况。本文提出 LILAC,这是一种全卷积的24 kHz 语音编解码器,工作频率为9.375 Hz,码率为0.75 kbit/s,其结构本身具备编解码幂等性:对任意有效 token 流解码得到的音频进行重编码,会返回完全相同的 token 流。LILAC 在实现幂等性的同时保持了竞争力的质量,在 LibriSpeech 和 LibriTTS-R 测试集上分别达到 UTMOS 4.14 和 4.24,与 sub-1 kbit/s 神经音频编解码器的当前最优水平(SOTA)相当。

英文摘要

Neural Audio Codecs are widely adopted in speech generation and editing. However, existing neural audio codecs are not idempotent: across the paper's twelve baseline systems, every configuration tested rewrites, on average, at least 15% of its tokens in a single decode-re-encode pass. This poses a problem for utilizing Neural Audio Codecs as token interfaces in pipelines where re-encoding decoded outputs can occur. We present LILAC, a fully convolutional 24 kHz speech codec at 9.375 Hz and 0.75 kbit/s that is codec idempotent by construction; re-encoding the decoded audio of any valid token stream returns the identical stream. LILAC achieves idempotency while maintaining competitive quality, reaching UTMOS 4.14 and 4.24 on LibriSpeech and LibriTTS-R test sets, comparable to SOTA sub-1 kbit/s Neural Audio Codecs.

Comments22 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑