arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AVE-Compass:迈向音频-视频编辑能力的整体评估

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu, Yibin Lin, Yikai Gu, Jialu Chen, Yuanxing Zhang, Jiaheng Liu

arXiv 2607.24821首次发表:更新:

发表机构

Nanjing University; Kuaishou Technology; National University of Singapore; Beijing University of Posts and Telecommunications; University of Illinois Urbana-Champaign(南京大学; 快手科技; 新加坡国立大学; 北京邮电大学; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有视频编辑基准未充分考虑视听耦合问题,提出AVE-Compass基准及AVE-Agent框架,通过多种评估方式评估编辑能力,能分解复杂指令迭代改进结果,提升了跨模态编辑等方面表现及感知质量。

AI 中文摘要

虽然基于指令的视频编辑发展迅速,但现实世界视频中的音频和视觉信号紧密耦合,编辑一种模态通常需要另一种模态的协调变化。现有基准主要评估无声剪辑上的视觉变换或孤立音频编辑,复杂的视听编辑和跨模态一致性未得到充分探索。我们引入AVE-Compass,一个具有145个精心策划的源视频、196个视听耦合编辑指令和2688个细粒度检查清单项目的综合基准。它通过基于检查清单的MLLM判断和专用的真实感评分标准评估指令遵循、保真度保持、真实感和编辑意图,并辅以自动跨模态、视频和音频指标。广泛评估表明,现有模型在执行跨模态指令时仍难以保留非目标内容。我们进一步提出AVE-Agent,一个模块化代理框架,将复杂指令分解为相关子任务,并通过自我反思和评估反馈迭代改进编辑结果。AVE-Agent在联合编辑中提高了指令执行、保真度保持和视听对齐,同时保持了具有竞争力的感知质量。

英文摘要

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑