发表机构
The Chinese University of Hong Kong, Shenzhen; Amphion Technology Co., Ltd.; Zhejiang University(香港中文大学(深圳); Amphion科技有限公司; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过对比预测边界与语言及声学参考,发现动态帧率编解码器边界含义取决于底层表示,且音节级边界对语义丰富的语音重建更有效。
AI 中文摘要
动态帧率神经语音编解码器用可变时长令牌取代均匀帧网格,使边界放置成为表示本身的一部分。然而,这些边界编码了什么,以及可解释的边界是否对神经语音重建有用,尚不清楚。本研究通过将预测边界与语言学和声学参考进行比较,结合边界解释和受控重建分析。我们发现边界含义取决于底层语音表示。对于面向ASR的SenseVoice和Whisper编码器,浅层强调语音、浊音和声学转换,而深层则转向音节和子词结构。在六种动态帧率算法和一种均匀(固定帧率)基线的比较中,更高级的语言对齐与更低的池化失真和更好的语义令牌重建相关。帧率匹配的边界测试使区别具体化:音节派生的分区优于均匀和相似性分区,而音素派生的分区并不优于均匀分区。我们推断,对于语义丰富的语音表示,有用的编解码器边界最好理解为围绕音节尺度组织的分配决策。
英文摘要
Dynamic-frame-rate neural speech codecs replace a uniform frame grid with variable-duration tokens, making boundary placement part of the representation itself. Yet it is unclear what these boundaries encode and whether interpretable boundaries are also useful for neural speech reconstruction. This work combines boundary interpretation and controlled reconstruction analysis by comparing predicted boundaries with linguistic and acoustic references. We find that boundary meaning depends on the underlying speech representation. For the ASR-oriented SenseVoice and Whisper encoders, shallow layers emphasize phonetic, voicing, and acoustic transitions, whereas deeper layers shift toward syllable and subword structure. In a comparison of six dynamic-frame-rate algorithms and a uniform (fixed-frame-rate) baseline, higher-level linguistic alignment is associated with lower pooling distortion and better reconstruction from semantic tokens. Frame-rate-matched boundary tests make the distinction concrete: a syllable-derived partition improves over both Uniform and Similarity, while a phoneme-derived partition does not improve over Uniform. We infer that for semantically rich speech representations, useful codec boundaries are best understood as allocation decisions organized around the syllable scale.
Comments5pages, 3 figures