arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

全双工语音模型的声学到文本KV压缩

Acoustic-to-Text KV Compression for Full-Duplex Speech Models

Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim

arXiv 2609.31224首次发表:更新:

发表机构

Sungkyunkwan University(成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对全双工语音模型内存密集问题,提出声学到文本KV压缩,利用聆听空隙将语音转文本并驱逐旧声学状态,减少64.6%峰值缓存,提升转录与问答性能。

AI 中文摘要

全双工语音语言模型持续累积声学键值(KV)状态,使得长时间运行的交互对内存的需求很高。在聆听过程中,模型可以在下一个音频单元到达之前完成对当前音频单元的处理;我们将此剩余时间间隔称为聆听时间空隙。我们提出声学到文本的KV压缩方法,该方法引入一个转录侧通道,在此时间间隔内将传入的语音转换为紧凑的文本记忆。当推理过程中缓存超过目标预算时,较旧的声学状态被驱逐,而转录文本和最近的声学上下文得以保留。我们使用LoRA对转录片段进行交叉熵训练来训练侧通道。为了保持聆听和说话行为,我们在原始模型的自然预测位置对其词元级输出分布应用知识蒸馏。在十分钟的LongSpeech会话中,我们的MiniCPM-o 4.5实现相比未进行驱逐的同一模型,将峰值流式KV缓存大小减少了64.6%。所提出的方法还在转录、时间问答和摘要任务上优于基线。Full-Duplex-Bench评估进一步显示了相当的停顿处理、轮流说话和打断性能。

英文摘要

Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inference, older acoustic states are evicted while transcripts and recent acoustic context remain. We train the side channel with LoRA using cross-entropy on transcription segments. To preserve listening and speaking behavior, we apply knowledge distillation to the original model's token-level output distributions at native prediction positions. On ten-minute LongSpeech sessions, our MiniCPM-o 4.5 implementation reduces peak streaming KV-cache size by 64.6% compared with the same model without eviction. The proposed method also improves transcription, temporal question answering, and summarization over the baseline. Full-Duplex-Bench evaluations further show comparable pause-handling, turn-taking, and interruption performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑