arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29121cs.CVcs.AI

少即是多:仅编码器的音视频分割

Less is More: Encoder-only Audio-Visual Segmentation

发表机构坦佩雷大学
查看机构详情
  • Tampere University(坦佩雷大学)

机构由 AI 辅助整理,请以论文原文为准。

Ilpo Viertola, Vladimir Iashin, Sophie Tötterström, Esa Rahtu

首次发表
浏览论文内容

中文总结 AI 辅助

提出仅编码器的音视频分割模型EASE,简化架构,实现365 FPS高速运行,训练不足11 GPU小时,在多种骨干和分辨率下达到最先进性能,证明AVSS可更简单高效。

中文摘要 AI 辅助

音视频语义分割(AVSS)旨在识别、分割并分类视频帧中发出声音的物体。以往的基于Transformer的AVSS方法在很大程度上继承了图像分割模型的设计原则。近期研究表明,这些图像分割模型包含对分割性能贡献甚微的冗余组件。基于这一见解,我们提出了仅编码器的音视频分割(EASE)。EASE的运行速度高达每秒365帧(FPS),在精度相当的情况下,比先前最先进的(SotA)AVS模型快3倍,并且训练时间不到11个GPU小时。此外,我们在不同骨干网络和输入分辨率下均取得了最先进的AVSS性能。我们的结果表明,AVSS可以更简单且更快,为未来研究和实时应用提供了可扩展的基础。代码、模型权重和样本可在以下网址获取:https URL

英文摘要

Audio-Visual Semantic Segmentation (AVSS) aims to identify, segment, and classify sound-emitting objects in video frames. Previous Transformer-based AVSS approaches largely inherit design principles from image segmentation models. Recent studies show that these image segmentation models contain redundant components that contribute little to the segmentation performance. Following this insight, we propose Encoder-only Audio-Visual Segmentation (EASE). EASE runs at up to 365 FPS, 3x faster than prior State-of-the-Art (SotA) AVS models at comparable accuracy, and trains in under 11 GPU-hours. Furthermore, we achieve SotA AVSS performance across different backbones and input resolutions. Our results demonstrate that AVSS can be both simpler and faster, providing a scalable foundation for future research and real-time applications. Code, model weights, and samples are available at https://ease-avs.notion.site

补充信息

↑