arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19523eess.AS

DAVSS:蒸馏视听状态空间模型

DAVSS: Distilled Audio-Visual State Space Models

Saurabhchand Bhati, Mrudula Athi, Amit S. Chhetri, James Glass

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出14M参数的DAVSS模型,通过小patch提升输入分辨率、30%参数用于联合建模,在体积远小于CAV-MAE等Transformer视听模型的同时,性能仍优于后者。

中文摘要 AI 辅助

由Transformer教师模型蒸馏得到的状态空间模型(SSMs)兼具Transformer的性能与SSMs的效率。我们将Transformer-SSM知识蒸馏扩展至多模态场景,提出蒸馏视听状态空间(DAVSS)模型。该模型参数规模为14M,较CAV-MAE等基于Transformer的模型小12倍,却仍能超越它们。DAVSS从两方面改进现有视听模型:1)更精细的输入分辨率:采用更小的patch尺寸处理输入,通过增加输入序列长度弥补模型规模的缩小,这一设计源于大patch尺寸会导致性能更低的观察结果;2)更深的联合建模:将模型的30%用于视听联合处理,而CAV-MAE中该比例不足5%,可在不显著增加视听拼接token计算成本的前提下实现更深的跨模态交互。

英文摘要

State-space models (SSMs) distilled from transformer teachers combine the performance of transformers with the efficiency of SSMs. We extend the Transformer-SSM knowledge distillation to a multimodal setting and propose the Distilled Audio-visual State-Space (DAVSS) model. The DAVSS model, 14M parameters, is 12 times smaller compared to transformer-based models such as CAV-MAE, and still outperforms them. DAVSS improves over the existing audio-visual models by: 1) Finer input resolution: using smaller patch sizes process the input, compensating for the smaller model size by increasing input sequence lengths. This is supported by the observation that a larger patch size results in lower performance. 2) Deeper joint modeling: utilizing a larger portion of the model (30%) for joint audio-visual processing, compared to <5% in CAV-MAE, enabling deeper cross-modal interaction without significantly increasing the computational cost associated with the concatenated audio-visual tokens.

↑