arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

听而能视:面向视听实例分割的状态感知聆听

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation

Leiye Liu, Miao Zhang, Jiahong Jiang, Jingjing Li, Jialong Zhong, Kai Peng, Tingwei Liu, Wei Ji, Yongri Piao, Huchuan Lu

arXiv 2608.03264首次发表:更新:

发表机构

Dalian University of Technology; Carnegie Mellon University; Yale University(大连理工大学; 卡内基梅隆大学; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 Hear to See(H2S)模型,通过 ASP 和 ADM 机制解决视听实例分割中重叠声源匹配与异步动态跟踪问题,在 AVISeg 数据集上实现 SOTA 性能,mAP 达 48.54,较此前方法提升 7.8%。

AI 中文摘要

视听实例分割(AVIS)需要以像素级掩码准确识别并跟踪单个发声对象。现有方法难以匹配重叠声学事件与视觉实例,且无法处理异步视听动态。因此产生两个关键问题:模型如何在重叠声源与视觉实例间建立精确对应,以及当视听信号时间错位时,模型如何维持鲁棒跟踪?本文提出 Hear to See(H2S)以应对这些挑战,通过两种机制实现:声学语义投影器(ASP)解耦混合音频,并建立从语义域到空间域的层级对应;异步动态调制器(ADM)通过音频调制的 Mamba 自适应调整状态转换,在动态变化时优先考虑当前信息,并在稳定时维持连续性。在 AVISeg 数据集上的实验表明,H2S 实现了 SOTA 性能,采用 COCO 预训练的 ResNet50 达到 48.54 mAP,较之前方法提升了 7.8%。论文被接受后代码将开源,源代码将在指定网址公开。

英文摘要

Audio-visual instance segmentation (AVIS) requires accurately identifying and tracking individual sounding objects with pixel-level masks. Existing methods struggle to match overlapping acoustic events with visual instances and handle asynchronous audio-visual dynamics. Therefore, two critical questions arise: how can a model establish precise correspondence between overlapping sound sources and visual instances, and how can a model maintain robust tracking when audio and visual signals are temporally misaligned?This paper proposes Hear to See (H2S), addressing these challenges through two mechanisms. The Acoustic-Semantic Projector (ASP) disentangles mixed audio and establishes hierarchical correspondence from semantic to spatial domains. The Asynchronous Dynamics Modulator (ADM) adaptively adjusts state transitions via audio-modulated Mamba, prioritizing current information during dynamic variations and maintaining continuity in stable periods.Experiments on AVISeg show H2S achieves SOTA performance, attaining 48.54 mAP with a COCO pretrained ResNet50 and surpassing the previous by 7.8\%. The code will be open-sourced once the paper is accepted. The source code will be publicly available at https://github.com/leiyeliu/H2S.

CommentsAccepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑