X-AuT:基于跨尺度蒸馏的语音大语言模型渐进式音频编码器压缩
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
浏览论文内容
中文总结 AI 辅助
X-AuT通过渐进式层剪枝和跨尺度蒸馏压缩语音LLM的音频编码器,在Qwen3-ASR上以更少参数降低错误率,优于直接剪枝。
中文摘要 AI 辅助
减少音频编码器的深度可以降低语音大语言模型的推理成本,但移除完整的块会扰动解码器所消耗的嵌入,并可能导致删除和过早的序列结束错误。我们引入了X-AuT,这是一个渐进式框架,通过简短的行为探针选择层组合,并通过表示对齐、跨尺度蒸馏、计划的学生策略监督和LoRA微调来恢复修剪后的模型。语言模型主干保持冻结,而注意力LoRA适配器和绑定的输出嵌入在蒸馏过程中进行调整。训练使用来自转录一致性流水线的最高一致性层级,随后在微调期间进行源重新加权。在十个公开的中英 benchmarks 上,将 Qwen3-ASR-0.6B 的音频编码器层从18层压缩到16层,将宏观平均错误率从5.61%降低到5.27%。14层模型达到5.75%的错误率,同时音频塔参数减少20.7%。在匹配的配方下,1.7B教师模型产生5.55%的平均错误率,而自蒸馏为8.45%,渐进式18→14层修剪优于直接修剪(5.75%对比6.73%)。这些单次运行结果确立了两个实用的操作点,并表明准确性效果因基准而异。项目网站:此 https URL
英文摘要
Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during distillation. Training uses the highest-agreement tier from a transcript-consistency pipeline, followed by source reweighting during finetuning. On ten public Chinese--English benchmarks, compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers reduces macro-average error from 5.61% to 5.27%. The 14-layer model reaches 5.75% with 20.7% fewer audio-tower parameters. Under the matched recipe, the 1.7B teacher yields 5.55% mean error, compared with 8.45% for self-distillation, and progressive 18$\rightarrow$14 pruning outperforms direct pruning (5.75% vs. 6.73%). These single-run results establish two practical operating points and show that the accuracy effects vary across benchmarks. Project website: https://xpeng-ai.github.io/x-aut
发表机构
- XPeng Inc.(小鹏汽车)
机构由 AI 辅助整理,请以论文原文为准。