AV-JEPA:将LeJEPA扩展到视听自监督学习
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
浏览论文内容
中文总结 AI 辅助
研究将LeJEPA扩展到视听自监督学习,核心方法是用早期融合视觉Transformer和模态丢弃作掩码训练模型对齐视图嵌入,主要贡献是实现跨模态对齐,架构简洁,在分类任务中有竞争力且支持零样本音频-视频检索。
中文摘要 AI 辅助
我们提出了AV-JEPA,这是LeJEPA到视听自监督学习的一个优雅的多模态扩展。使用早期融合视觉Transformer和模态丢弃作为掩码,训练该模型对齐全局和各模态局部视图的嵌入,同时SIGReg目标鼓励理论上最优的分布。这在潜在空间中实现了跨模态对齐,得到了一个无解码器、无EMA教师、无复杂多术语损失或对比负样本的简洁架构。所提出的AV-JEPA主干在VGGSound上实现了有竞争力的分类性能(top-1为57.1%)和AudioSet上(32.7 mAP),并支持开箱即用的零样本音频-视频检索。
英文摘要
We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of global and per-modality local views, while the SIGReg objective encourages a theoretically optimal distribution. This achieves cross-modal alignment in the latent space, resulting in a remarkably clean architecture with no decoder, EMA teacher, complex multi-term losses, or contrastive negatives. The proposed AV-JEPA backbone delivers competitive classification performance on VGGSound (57.1% top-1) and AudioSet (32.7 mAP) and supports zero-shot audio-video retrieval out of the box.
发表机构
- ELLIS Institute Finland(芬兰ELLIS研究所)
- Department of Computer Science, Aalto University, Espoo, Finland(艾尔沃斯大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。