发表机构
Nankai University; Tencent(南开大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对多指令编辑长视频的挑战,提出MMLVE任务与智能体编辑框架,构建MMLVE-Bench数据集及评估指标,所提MMLVE-Agent优于Seedance 2.0等闭源SOTA方法,可实现一致且无幻觉的长视频编辑。
AI 中文摘要
尽管生成式AI已显著推动视频编辑发展,但现有方法主要聚焦于单镜头或短视频片段,用多条指令编辑长视频仍是巨大挑战。固定时长分割等朴素分块策略常导致实体碎片化、严重编辑幻觉及时空连续性中断。为弥合这一差距,本文提出多指令多镜头长视频编辑(Multi-Instruction Multi-Shot Long-Video Editing, MMLVE)任务,围绕三个核心目标构建:跨镜头编辑一致性(Cross-Shot Editing Consistency, CSEC)、多指令解耦(Multi-Instruction Decoupling, MID)、时空结构零破坏(Zero-Destruction on Spatiotemporal Structure, ZDSS)。为应对这三个独特挑战,本文引入智能体编辑框架,利用大语言模型(Large Language Models, LLMs)与视觉语言模型(Vision-Language Models, VLMs)的协同实现镜头级视频解耦与精准指令解析。此外,为全面评估该任务,本文构建MMLVE-Bench,这是一个聚焦MMLVE的数据集,具有复杂的现实世界时空动态、高密度异构指令及稀疏随机实体分布的特征。本文还进一步开发了三个聚焦MMLVE的评估指标以评估编辑结果质量。大量实验表明,本文的MMLVE-Agent优于现有闭源SOTA方法(如Seedance 2.0),成功消除编辑幻觉、保持跨镜头编辑一致性并实现无缝时空过渡。
英文摘要
While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or short video clips. Editing long videos with multiple instructions remains a formidable challenge. Naive chunking strategies, e.g., fixed-duration segmentation, often lead to entity fragmentation, severe editing hallucinations, and disrupted temporal continuity. To bridge this gap, we introduce the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, which is structured around three core objectives: Cross-Shot Editing Consistency (CSEC), Multi-Instruction Decoupling (MID), and Zero-Destruction on Spatiotemporal Structure (ZDSS). To tackle these three unique challenges, we introduce an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing. Furthermore, to comprehensively evaluate this task, we construct MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions. Three MMLVE-focused evaluation metrics are further exploited to assess the quality of the editing results. Extensive experiments demonstrate that our MMLVE-Agent outperforms existing closed-source SOTA approaches (e.g., Seedance 2.0), successfully eliminating editing hallucinations, preserving cross-shot editing consistency, and attaining seamless spatiotemporal transitions.
CommentsProject Page: https://wucy0519.github.io/MMLVE/ and see source codes at https://github.com/Wucy0519/MMLVE