发表机构
ELLIS Institute Finland; Aalto University(芬兰ELLIS研究所; 阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LeAVJEPA提出一种极简视听自监督架构,利用单一早期融合ViT和模态丢弃实现跨模态对齐,在AudioSet-20K、ESC-50和VGGSound上取得优异性能,并支持零样本检索。
AI 中文摘要
先前的视听自监督学习方法依赖于EMA目标编码器、预测头、重建解码器和对比损失等机制。我们引入了LeAVJEPA,这是第一个在LeJEPA的无崩溃目标下训练的视听编码器。一个单一的早期融合视觉Transformer处理音频、视频和联合音视频输入。模态丢弃将缺失的模态视为同一事件的另一个视图,使跨模态对齐隐含在目标中。该模型将全局嵌入与模态特定的局部嵌入对齐,SIGReg防止表示崩溃。一项受控消融实验确定模态丢弃是视听对齐的关键机制。尽管架构简单,LeAVJEPA在冻结评估下于AudioSet-20K上达到36.0 mAP,在ESC-50上达到91.3%的准确率。微调后,它在VGGSound上达到61.1%的准确率,其嵌入支持零样本视听检索。
英文摘要
Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes audio, video, and joint audio-video inputs. Modality dropout treats a missing modality as another view of the same event, making cross-modal alignment implicit in the objective. The model aligns global embeddings with modality-specific local embeddings, and SIGReg prevents representational collapse. A controlled ablation identifies modality dropout as the key mechanism for audio-visual alignment. Despite the architectural simplicity, LeAVJEPA reaches 36.0 mAP on AudioSet-20K and 91.3% accuracy on ESC-50 under frozen evaluation. After fine-tuning, it reaches 61.1% accuracy on VGGSound, and its embeddings support zero-shot audio-visual retrieval.