arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AVENUE:音频-视频编辑的理解与评估

AVENUE: Audio-Video EditiNg Understanding and Evaluation

Hayeon Kim, Yoojin Jang, Jaejun Yoo

arXiv 2609.04253首次发表:更新:

发表机构

Ulsan National Institute of Science and Technology (UNIST)(蔚山国立科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出AVENUE基准与模态感知评估框架,分析三类AV编辑范式的模态选择性,发现现有模型编辑单模态时易对另一模态产生非预期修改,为可控AV编辑模型发展提供支撑。

AI 中文摘要

音频-视频(AV)编辑旨在根据目标提示修改音频和视频内容。与单模态编辑不同,AV编辑要求模型仅从提示中推断出模态选择性编辑范围:不仅要确定应修改的内容,还要确定应保留的模态。因此,要对这类模型进行忠实评估,需要满足两个条件:一是涵盖多种编辑类型和模态类别的基准,二是本身具备模态感知和样本特异性的评估方法。然而,现有的AV编辑基准对编辑类型和模态组合的覆盖有限,而当前的评估系统通常是模态盲的、样本无关的,难以评估模型是否忠实地保留了非目标模态。为解决这些不足,我们推出了AVENUE(Audio-Video EditiNg Understanding and Evaluation),包含两项贡献:(1)一个基准,包含1291个源片段和7957条编辑指令,涵盖音频目标型、视频目标型以及AV耦合型编辑类型,从VGGSound中筛选并经人工验证;(2)一个样本特异性、模态感知的评估框架,为每个样本指定预期修改内容和必须保留的内容。我们评估了涵盖联合、序列、分离三种编辑范式的代表性AV编辑模型,首次对不同范式下的模态选择性进行了系统分析。研究发现了一个根本性的开放挑战:无论采用何种范式,现有模型在编辑某一模态时,频繁会对另一模态产生非预期修改。AVENUE提供了基准和模态感知评估框架,以推动更可控的AV编辑模型的发展。我们的数据集在Hugging Face上公开可用:this https URL。

英文摘要

Audio-video (AV) editing aims to modify audio and video content according to a target prompt. Unlike single-modality editing, AV editing requires models to infer a modality-selective edit scope from the prompt alone: determining not only what should change, but also which modality should be preserved. Faithfully evaluating such models therefore requires both (i) benchmarks that span diverse edit types and modality categories, and (ii) evaluation that is itself modality-aware and sample-specific. However, existing AV editing benchmarks provide limited coverage of edit types and modality combinations, while current evaluation systems are often modality-blind and sample-agnostic, making it difficult to assess whether models faithfully preserve the unintended modality. To address these gaps, we introduce AVENUE, Audio-Video EditiNg Understanding and Evaluation, comprising two contributions: (1) a benchmark of 1,291 source clips and 7,957 editing instructions across audio-targeted, video-targeted, and AV-coupled edit types, curated and human-verified from VGGSound; and (2) a sample-specific, modality-aware evaluation framework that specifies, for each sample, both the intended change and the content that must remain intact. We evaluate representative AV editing models spanning three editing paradigms : joint, sequential, and separate, providing the first systematic analysis of modality-selectivity across paradigms. Our findings reveal a fundamental open challenge: when editing one modality, existing models frequently induce unintended changes in the other, regardless of paradigm. AVENUE provides a benchmark and modality-aware evaluation framework to drive progress toward more controllable AV editing models. Our dataset is publicly available on Hugging Face: https://huggingface.co/datasets/AVENUE-dataset/AVENUE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑