发表机构
Dalian University of Technology; Carnegie Mellon University; Yale University(大连理工大学; 卡内基梅隆大学; 耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 Hear to See(H2S)模型,通过 ASP 和 ADM 机制解决视听实例分割中重叠声源匹配与异步动态跟踪问题,在 AVISeg 数据集上实现 SOTA 性能,mAP 达 48.54,较此前方法提升 7.8%。
AI 中文摘要
视听实例分割(AVIS)需要以像素级掩码准确识别并跟踪单个发声对象。现有方法难以匹配重叠声学事件与视觉实例,且无法处理异步视听动态。因此产生两个关键问题:模型如何在重叠声源与视觉实例间建立精确对应,以及当视听信号时间错位时,模型如何维持鲁棒跟踪?本文提出 Hear to See(H2S)以应对这些挑战,通过两种机制实现:声学语义投影器(ASP)解耦混合音频,并建立从语义域到空间域的层级对应;异步动态调制器(ADM)通过音频调制的 Mamba 自适应调整状态转换,在动态变化时优先考虑当前信息,并在稳定时维持连续性。在 AVISeg 数据集上的实验表明,H2S 实现了 SOTA 性能,采用 COCO 预训练的 ResNet50 达到 48.54 mAP,较之前方法提升了 7.8%。论文被接受后代码将开源,源代码将在指定网址公开。
英文摘要
Audio-visual instance segmentation (AVIS) requires accurately identifying and tracking individual sounding objects with pixel-level masks. Existing methods struggle to match overlapping acoustic events with visual instances and handle asynchronous audio-visual dynamics. Therefore, two critical questions arise: how can a model establish precise correspondence between overlapping sound sources and visual instances, and how can a model maintain robust tracking when audio and visual signals are temporally misaligned?This paper proposes Hear to See (H2S), addressing these challenges through two mechanisms. The Acoustic-Semantic Projector (ASP) disentangles mixed audio and establishes hierarchical correspondence from semantic to spatial domains. The Asynchronous Dynamics Modulator (ADM) adaptively adjusts state transitions via audio-modulated Mamba, prioritizing current information during dynamic variations and maintaining continuity in stable periods.Experiments on AVISeg show H2S achieves SOTA performance, attaining 48.54 mAP with a COCO pretrained ResNet50 and surpassing the previous by 7.8\%. The code will be open-sourced once the paper is accepted. The source code will be publicly available at https://github.com/leiyeliu/H2S.
CommentsAccepted by ACM MM 2026