arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08569cs.AIcs.SD

VoxZip:面向长上下文音频推理的语义锚定时序KV缓存压缩

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference

  • Zhejiang University(浙江大学)
  • Meituan(美团)

机构由 AI 辅助整理,请以论文原文为准。

Wenxu Jia, Dongjie Fu, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin

AI总结:

针对语音大语言模型长上下文推理的KV缓存内存瓶颈,提出无训练两阶段语义锚定框架VoxZip,在多音频基准上实现高效压缩,维持高性能并提升推理效率。

AI中文摘要:

近期,语音大语言模型的进展已展现出理解复杂音频任务的卓越能力。尽管取得了这些进展,但其长上下文推理仍受到过高的KV缓存内存需求的严重瓶颈。现有的以文本为中心的压缩方法在此处表现不佳,往往会破坏语音连续性或丢弃关键语义线索。为解决这一问题,我们提出了VoxZip,一种无训练的两阶段语义锚定KV缓存压缩框架。第一阶段使用自动语音识别(ASR)转录文本作为显式语义锚点,对音频令牌进行时序对齐、压缩和融合,显著降低初始KV缓存的同时提升令牌信息密度。为进一步提高压缩率,第二阶段采用基于时序衰减累积注意力的动态过滤策略,在驱逐非必要令牌的同时缓解早期令牌偏差。在Qwen3-Omni上针对六个不同音频基准的综合评估表明,我们的方法具有优越性。VoxZip在长音频推理中表现出色,并在短音频任务中始终保持高保真感知。值得注意的是,即使在长上下文场景下采用激进的20倍KV缓存压缩,它仍能维持超过90%的未压缩基线性能。此外,在4倍压缩率下,VoxZip的推理吞吐量提高了1.9倍,同时峰值内存开销降低了3.3倍。代码和模型将在此https URL上提供。

英文摘要:

Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks. Despite this progress, their long-context inference remains severely bottlenecked by prohibitive KV cache memory demands. Existing text-centric compression methods struggle here, often disrupting speech continuity or discarding crucial semantic cues. To address this, we propose VoxZip, a train-free, two-stage semantic-anchored KV cache compression framework. The first stage uses automatic speech recognition (ASR) transcriptions as explicit semantic anchors to temporally align, compress, and fuse audio tokens, significantly reducing the initial KV cache while elevating token information density. To further improve the compression ratio, the second stage employs a dynamic filtering strategy based on temporally decayed accumulated attention to evict non-essential tokens while mitigating early-token bias. Comprehensive evaluations on Qwen3-Omni across six diverse audio benchmarks demonstrate the superiority of our approach. VoxZip excels in long-audio reasoning and consistently maintains high-fidelity perception on short-form tasks. Notably, it sustains over 90\% of the uncompressed baseline performance even under an aggressive 20x KV cache compression in long-context scenarios. Furthermore, at a 4x compression ratio, VoxZip yields a 1.9x increase in inference throughput alongside a 3.3x reduction in peak memory overhead. Code and models will be available at https://github.com/MM-Speech/VoxZip.

补充信息

↑