arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10659cs.SD

DINO-A:将自蒸馏视觉Transformer适配至通用音频表示学习

DINO-A: Adapting Self-Distillation Vision Transformers to General Audio Representation Learning

  • Warsaw University of Technology(华沙理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Tomasz Radzikowski, Mateusz Modrzejewski, Przemysław Rokita

AI总结:

本研究提出DINO-A,将自蒸馏视觉Transformer适配至通用音频表示学习,通过替换输入模态和增强方式实现,在多音频数据集上验证了其性能,发现了补丁分辨率、骨干选择对任务的影响及与BYOL-A v2的性能差异原因。

AI中文摘要:

我们提出了DINO-A,这是一种将自蒸馏方法从视觉领域适配至通用音频表示学习的技术。尽管DINO已成为自监督视觉学习中的经典方法,且现有音频领域的研究已探索了潜在预测(BYOL-A)和掩码建模(Audio-MAE、BEATs),但尚无研究像BYOL-A适配BYOL那样将经典DINO方法应用于通用音频分类。DINO-A保留了DINO的多裁剪、EMA教师模型和高维投影结构,仅将输入模态和增强方式替换为对数梅尔频谱图和BYOL-A v2增强模块。我们在FSD50K数据集上对三种骨干网络进行预训练,包括两个采用8×8和16×16补丁的Vision Transformer及一个卷积编码器,并在ESC-50、Speech Commands v2、UrbanSound8K和GTZAN数据集上通过线性探测进行评估。研究得出三项关键发现:Vision Transformer家族内的补丁分辨率对表示质量有一致影响,较小的补丁在所有四项任务中表现更优;Vision Transformer与卷积骨干的选择与任务类型存在交互,卷积网络在语音任务中表现领先,而Vision Transformer在环境声音和音乐任务中表现领先;在相同预训练和评估条件下,DINO-A与BYOL-A v2的平均性能差异为11.96个百分点,我们将该差异归因于两种机制:DINO的高维投影空间与FSD50K有限规模的交互,以及DINO采用而BYOL-A v2未采用的多裁剪增强带来的额外成本,而在视觉领域中至关重要的高维投影空间,在FSD50K规模下反而成为劣势。

英文摘要:

We present DINO-A, an adaptation of self-distillation from vision to general audio representation learning. While DINO has become a canonical method in self-supervised vision and prior audio work has explored latent prediction (BYOL-A) and masked modeling (Audio-MAE, BEATs), no prior work has brought canonical DINO to general audio classification in the way BYOL-A brought BYOL. DINO-A retains DINO's multi-crop, EMA teacher, and high-dimensional projection, replacing only the input modality and augmentations with log-mel spectrograms and the BYOL-A v2 augmentation block. We pretrain three backbones, two Vision Transformers with 8x8 and 16x16 patches and a convolutional encoder, on FSD50K and evaluate them with linear probing on ESC-50, Speech Commands v2, UrbanSound8K, and GTZAN. Three findings characterize the resulting representations. Patch resolution within the Vision Transformer family has consistent effect on representation quality, with smaller patches winning across all four tasks. The choice between Vision Transformer and convolutional backbone interacts with task type: convolutional networks lead on speech while Vision Transformers lead on environmental sounds and music. Under identical pretraining and evaluation conditions, DINO-A and BYOL-A v2 differ by 11.96 percentage points on average, and we trace this difference to two mechanisms: the interaction between DINO's high-dimensional projection space and FSD50K's limited scale, and the additional cost of multi-crop augmentation, which DINO uses but BYOL-A v2 does not. The high-dimensional projection space, central to DINO's success in vision, becomes a liability at FSD50K scale.

↑