arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01492cs.SDcs.CLeess.AS

Q-SPT:面向低帧率语音分词的可学习查询式压缩

Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization

Jeeyoung Yun, Seohwan Yun, Sungwoong Kim

首次发表
浏览论文内容

中文总结 AI 辅助

Q-SPT提出可学习查询式双流压缩器,分别处理语义和声学流,并辅以文本损失监督,在低帧率下实现最优重建、识别与合成质量。

中文摘要 AI 辅助

神经语音编解码器日益充当语音语言模型(SLM)的分词器。降低帧率可减少SLM的计算和内存成本,但难以同时保留语言信息和声学细节。现有方法依赖基于规则的压缩:平均池化可能丢弃语言信息,而基于相似度的合并则对相邻帧相似度使用固定阈值,并将所得边界应用于声学流。我们提出Q-SPT,一种低帧率双流语音分词器,其具有针对语义和声学表示分别优化的、上下文感知的、可学习的查询式压缩器。具体而言,固定速率的查询独立地将语义流和声学流作为不同的键值源进行关注,通过两个分别学习的压缩器实现流特定的、上下文感知的聚合。此外,自回归文本损失显式监督语义压缩器以保留语言信息。实验结果表明,在相同帧率下,Q-SPT在所评估的编解码器中实现了最佳重建质量。在下游SLM中,它取得了最佳语音识别准确率和文本到语音感知质量,且可懂度具有竞争力。

英文摘要

Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, whereas similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream. We propose Q-SPT, a low-frame-rate dual-stream speech tokenizer with separate, context-aware, learnable query-based compressors specialized for semantic and acoustic representations. In particular, queries at a fixed rate independently attend to the semantic and acoustic streams as separate key-value sources, enabling stream-specific, context-aware aggregation through two separately learned compressors. In addition, an autoregressive text loss explicitly supervises the semantic compressor to preserve linguistic information. Experimental results show that Q-SPT achieves the best reconstruction among the evaluated codecs at the same frame rate. In downstream SLMs, it yields the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.

发表机构

  • Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

↑