arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36503cs.AI

AVIO:在视听场景中学习添加和移除发声物体

AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes

Weihan Xu, Kan Jen Cheng, Koichi Saito, Jingyu Shi, Tingle Li, Yisi Liu, Liming Wang, Masato Ishii, Takashi Shibuya, Gopala Anumanchipalli, Paul Pu Liang

首次发表
浏览论文内容

中文总结 AI 辅助

针对视听场景中发声物体添加与移除的成对监督数据稀缺问题,提出AVIOBench数据集和AVIO模型,通过源条件特征调制与参考帧课程实现联合编辑,实验验证了方法的有效性。

中文摘要 AI 辅助

添加或移除一个发声物体需要协调改变视觉内容和声音,同时保留周围场景。然而,针对局部非语音视听编辑的成对监督仍然有限,因为视觉和声学编辑必须针对同一物体,并将其声音从重叠声源中隔离出来。为解决这一差距,我们引入了\ extit{AVIOBench}数据集,该数据集包含37.9小时的成对视听示例,涵盖1,878个目标物体名称。AVIOBench通过共享身份和视觉掩码将每个目标物体的视觉存在和声学贡献联系起来。我们的自动化流程使用视觉定位和跨模态一致性来选择目标声音移除候选,然后联合优化视听对以提高感知质量和跨模态一致性。基于该数据集,我们提出了\ extit{AVIO},它通过源条件特征调制适配预训练的文本到视听生成模型,以联合学习物体添加和移除。参考帧课程在训练期间逐渐减少参考条件,使单一模型能够执行仅指令编辑,并可选视觉引导。定量和定性评估表明,视听物体移除和添加效果有效,可选的参考引导为添加提供了外观和位置控制。

英文摘要

Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.

发表机构

  • University of Washington(华盛顿大学)
  • Massachusetts Institute of Technology(麻省理工学院)
  • University of Maryland, College Park(马里兰大学帕克分校)
  • University of California, Berkeley(加州大学伯克利分校)
  • Sony Research(索尼研究)

机构由 AI 辅助整理,请以论文原文为准。

↑