arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用语义和视觉线索增强神经语音编码

Enhancing Neural Speech Coding with Semantic and Visual Cues

Yao Guo, Yang Ai, Hui-Peng Du, Xiao-Hang Jiang, Chen-Yuan Ning, Zhen-Hua Ling

arXiv 2609.05076首次发表:更新:

发表机构

National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China(中国科学技术大学语音和语言信息处理国家工程研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出SVSC编解码器,通过语义与视觉线索的两种注入策略,在低比特率下提升神经语音编码的重建性能,使ViSQOL评分达4.01

AI 中文摘要

在低比特率下,神经语音编解码器的容量有限,无法编码高质量重建所需的全部信息,尤其是在仅依赖语音衍生表示时。为解决这一局限,本文提出语义与视觉增强语音编解码器(SVSC),将语义和视觉线索融入神经语音编码过程。具体而言,SVSC构建于主流神经语音编码架构之上,引入语义编解码分支与图像分析合成分支;通过交叉注意力机制将深度语义特征与视觉线索融合,形成富含上下文和发音信息的辅助高层表示。为适配不同推理场景,SVSC基于辅助语义和视觉线索的可用性,引入两种信息注入策略:当存在此类线索时,融合模式通过特征拼接将辅助表示直接纳入语音编码分支;否则,蒸馏模式在训练期间通过知识蒸馏将辅助信息传递至语音编码分支,实现无额外输入的纯语音推理。实验结果验证了融入语义和视觉线索的有效性,使重建语音的ViSQOL评分从3.86提升至4.01。

英文摘要

At low bitrates, neural speech codecs have limited capacity to encode all information needed for high-quality re construction, especially when relying solely on speech-derived representations. To address this limitation, this paper proposes a Semantic- and Visual-enhanced Speech Codec (SVSC), which in corporates semantic and visual cues into the neural speech coding process. Specifically, built upon a mainstream neural speech cod ing architecture, SVSC introduces a semantic encoding-decoding branch and an image analysis-synthesis branch. It fuses deep semantic features with visual cues through a cross-attention mech anism, forming an auxiliary high-level representation enriched with contextual and articulatory information. To handle different inference scenarios, SVSC introduces two information-injection strategies based on the availability of auxiliary semantic and vi sual cues. When such cues are available, the fusion mode directly incorporates the auxiliary representations into the speech coding branch through feature concatenation; otherwise, the distillation mode transfers auxiliary information into the speech coding branch through knowledge distillation during training, enabling speech-only inference without additional inputs. Experimental results validate the effectiveness of incorporating semantic and visual cues, improving the ViSQOL score of reconstructed speech from 3.86 to 4.01.

Comments6 pages, 2 figures, submitted to APSIPA 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑