arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15295cs.MMcs.AIcs.LGcs.SD

AV-JEPA:将LeJEPA扩展到视听自监督学习

AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin

首次发表
浏览论文内容

中文总结 AI 辅助

研究将LeJEPA扩展到视听自监督学习,核心方法是用早期融合视觉Transformer和模态丢弃作掩码训练模型对齐视图嵌入,主要贡献是实现跨模态对齐,架构简洁,在分类任务中有竞争力且支持零样本音频-视频检索。

中文摘要 AI 辅助

我们提出了AV-JEPA,这是LeJEPA到视听自监督学习的一个优雅的多模态扩展。使用早期融合视觉Transformer和模态丢弃作为掩码,训练该模型对齐全局和各模态局部视图的嵌入,同时SIGReg目标鼓励理论上最优的分布。这在潜在空间中实现了跨模态对齐,得到了一个无解码器、无EMA教师、无复杂多术语损失或对比负样本的简洁架构。所提出的AV-JEPA主干在VGGSound上实现了有竞争力的分类性能(top-1为57.1%)和AudioSet上(32.7 mAP),并支持开箱即用的零样本音频-视频检索。

英文摘要

We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of global and per-modality local views, while the SIGReg objective encourages a theoretically optimal distribution. This achieves cross-modal alignment in the latent space, resulting in a remarkably clean architecture with no decoder, EMA teacher, complex multi-term losses, or contrastive negatives. The proposed AV-JEPA backbone delivers competitive classification performance on VGGSound (57.1% top-1) and AudioSet (32.7 mAP) and supports zero-shot audio-video retrieval out of the box.

发表机构

  • ELLIS Institute Finland(芬兰ELLIS研究所)
  • Department of Computer Science, Aalto University, Espoo, Finland(艾尔沃斯大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

↑