发表机构
Xiaomi Inc.(小米公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TV-AudioRemover,结合视觉编辑视频与文本指令,通过多任务训练和难混合课程实现目标声音移除,并在新基准上达到最先进性能。
AI 中文摘要
视觉对象移除可以从视频帧中消除目标,但其声学痕迹仍保留在音轨中,导致明显的视听不一致。现有的视频修复模型仅基于像素操作,而音频编辑模型,尤其是针对声音移除任务的模型,通常由文本驱动,因此依赖有限的单模态控制,其效果不如提供更强语义基础和时序同步线索的多模态引导。在本文中,我们提出了文本-视觉引导的声音移除(TV-AudioRemover),一种目标声音移除框架,利用视觉编辑后的视频以及自然语言指令,从原始音频混合中抑制与移除视觉对象相关的声音。为获取高质量训练数据,我们设计了一条流程来构建百万级规模的单对象视听对齐样本数据集,并从中合成针对模型训练定制的混合-目标对。为有效利用视觉上下文并遵循指令意图,我们通过任务令牌、可泛化的指令建模和模态特定的全局引导来增强模型架构。我们进一步采用多任务训练以强化任务角色理解,并在微调阶段使用难混合课程,利用语义相似的声学混合来增强细粒度源区分。为支持评估,我们提出了AV-Remove-Bench,一个全面的视听对象移除基准,以及专用的客观指标和基于MLLM的评估协议。实验表明,我们的方法在主观和客观指标上均达到了最先进的性能。项目页面:此https URL。
英文摘要
Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual inconsistency. Existing video inpainting models operate solely on pixels, while audio editing models, especially for the sound removal task, are typically driven by text and therefore rely on limited single-modal control, which is less effective than multimodal guidance that provides stronger semantic grounding and temporal synchronization cues. In this paper, we present Text-Visual Guided Sound Removal (TV-AudioRemover), a target sound removal framework that leverages the visually edited video together with a natural-language instruction to suppress the sound associated with the removed visual object from the original audio mixture. To acquire high-quality training data, we devise a pipeline to construct a million-scale dataset of single-object audio-visual aligned samples, from which we synthesize mixture-target pairs customized for model training. To effectively leverage visual context and follow instruction intent, we augment the model architecture with task tokens, generalizable instruction modeling, and modality-specific global guidance. We further adopt multi-task training to strengthen task-role comprehension, and employ a hard-mixture curriculum that leverages semantically similar acoustic mixtures during fine-tuning to enhance fine-grained source discrimination. To support evaluation, we present AV-Remove-Bench, a comprehensive audio-visual object removal benchmark, along with dedicated objective metrics and an MLLM-based evaluation protocol. Experiments demonstrate that our method achieves state-of-the-art performance on both subjective and objective metrics. Project page: https://yjx-research.github.io/TV-AudioRemover/.