arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16328cs.CV

GRNEdit:生成式细化网络中基于二元证据视角的高效通用视频编辑

GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks

Feng Xie, Jiagao Hu, Fuhao Li, Zepeng Wang, Yuxuan Chen, Dahua Gao, Fei Wang, Daiguo Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

GRNEdit是轻量级两阶段通用视频编辑框架,以二元证据视角建模编辑意图,仅用0.6M数据和少量参数训练,在OpenVE-Bench上表现优于或媲美同类开源编辑器。

中文摘要 AI 辅助

基于指令的通用视频编辑旨在通过单一直观界面整合多样的编辑操作。现有方法常依赖资源密集型条件设置,要么使用重量级分支,要么采用代价高昂的源拼接。是否存在高效的方式来建模编辑意图?为此,我们提出GRNEdit,这是一个轻量级两阶段框架。GRN通过位的组合编码视觉语义,为我们的方法提供了灵感。通过特定任务微调,我们进一步拓展该表示,将编辑语义重新定义为对单个位的局部保留或翻转决策。源信息因此被建模为支持观测到的二元状态的逐坐标证据,而GRN主干负责将这些证据的全局组合解析为连贯的生成语义。在第一阶段,一个紧凑的编码器将离散源代码转换为连续证据信号,GRN在整个二元细化过程中同化这些信号。受无分类器引导的空提示训练启发,我们进一步赋予空条件特定于编辑的含义:空指令表示无编辑,并通过源重建进行监督。该身份路径不仅在第一阶段隐含地增强了证据利用和内容保留,还在与编辑状态相同的表示空间中生成了源保留状态。因此,第二阶段可直接将每个编辑状态与其源保留对应物进行比较,并利用它们的差异来修正未解决的目标位决策。GRNEdit仅在0.6M对数据上进行训练,且条件参数占比不足3%,其中GRNEdit-2B和GRNEdit-8B在OpenVE-Bench上分别取得4.03和4.18的分数。2B模型的性能优于多个14B开源编辑器,而8B模型的表现与领先的开源编辑器相当。

英文摘要

Instruction-based general video editing seeks to unify diverse editing operations within a single, intuitive interface. Existing approaches often rely on resource-intensive conditioning, using either heavyweight branches or costly source concatenation. Is there any efficient way to model editing intent? Thus, we introduce GRNEdit, a lightweight two-stage framework. GRN inspires our approach by encoding visual semantics through combinations of bits. Through task-specific fine-tuning, we take this representation further and recast editing semantics as local retain-or-flip decisions over individual bits. Source information is consequently modeled as coordinate-wise evidence supporting the observed binary states, while the GRN backbone remains responsible for resolving their global composition into coherent generative semantics. In Stage I, a compact encoder translates discrete source codes into continuous evidence signals, which GRN assimilates throughout binary refinement. Inspired by null-prompt training for classifier-free guidance, we further assign the null condition an editing-specific meaning: an empty instruction denotes no edit and is supervised through source reconstruction. This identity pathway not only implicitly strengthens evidence utilization and content preservation in Stage I, but also produces a source-preserving state in the same representation space as the edited state. Stage II can therefore directly compare each edited state with its source-preserving counterpart and use their discrepancy to revise unresolved target-bit decisions. Trained on only 0.6M pairs with less than 3\% conditioning parameters, GRNEdit-2B and GRNEdit-8B achieve scores of 4.03 and 4.18 on OpenVE-Bench. The 2B model outperforms multiple 14B open-source editors, while the 8B model performs on par with leading open-source editors.

发表机构

  • Xidian University(西安电子科技大学)
  • Xiaomi Inc.(小米公司)

机构由 AI 辅助整理,请以论文原文为准。

↑