arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CARD:用于无编码器音频字幕的跨组件音频表示蒸馏

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning

Ganesh Pavan Kartikeya Bharadwaj Kolluri, Yuchen Zhang, Michael Kampouridis, Ravi Shekhar

arXiv 2607.04619首次发表:更新:

发表机构

School of Computer Science and Electronic Engineering, University of Essex; Institute for Analytics and Data Science, University of Essex(埃塞克斯大学计算机科学与电子工程学院; 埃塞克斯大学分析与数据科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出无编码器音频字幕模型CARD,去除推理时的编码器,用含合并LoRA适配器的投影仪给冻结的语言模型供能。通过跨组件蒸馏预训练音频教师模型,提升了CIDEr-D指标。

AI 中文摘要

现代自动音频字幕系统通过可训练投影仪将冻结音频编码器与大语言模型配对,带来编码器推理成本等问题。我们提出CARD,一种推理时无编码器的音频字幕模型,用含合并LoRA适配器的1320万参数投影仪给冻结大语言模型供能,训练教师模型被丢弃。CARD将预训练音频教师模型蒸馏到模型中,跨组件路由教师表征,提升了CIDEr-D指标。

英文摘要

Modern automated audio captioning systems pair a frozen audio encoder with a large language model (LLM) via a trainable projector, incurring the encoder's inference cost and bottlenecking the model through its fixed acoustic features. We present CARD, an encoder-free audio captioning model that removes the encoder at inference: a 13.2M projector feeds a frozen LLM with merged LoRA adapters, while the teacher used to train it is discarded. CARD distills a pre-trained audio teacher (CLAP-HTSAT) into the model, but rather than injecting it into the LLM alone, it routes the teacher's representations across components: perceptual stages to the projector and semantic stages to the LLM. This placement improves CIDEr-D by +11.9 over an LLM-only distilled model on AudioCaps and by +5.0 on Clotho, reaching 55.4 against a 66.4 encoder-kept upper bound with no encoder at inference, showing that where a teacher's knowledge is placed matters as much as its presence.

CommentsAccepted to IEEE Spoken Language Technology (SLT) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑