arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于视频对象插入和层分解的显式层建模

Explicit Layer Modeling for Video Object Insertion and Video Layer Decomposition

Kyujin Han, Seungjoo Shin, Sunghyun Cho

arXiv 2607.25802首次发表:更新:

发表机构

POSTECH(浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视频编辑系统缺乏显式分层表示的问题,提出TriLayer数据集及DBL-Diffusion框架,用于视频对象插入和层分解任务,显著提升了插入保真度和分解质量。

AI 中文摘要

大多数视频编辑系统仍缺乏显式分层视频表示,限制了其进行逼真合成、对象重用和一致操作的能力。在视频对象插入和视频层分解中这种限制尤为明显,现有方法因缺乏显式前景层监督而依赖隐式推理或逐场景优化。我们引入TriLayer数据集,在此基础上提出DBL-Diffusion框架,用于视频对象插入和层分解任务,实验表明显式层建模显著提高了插入保真度和分解质量。

英文摘要

Most video editing systems still lack explicit layered video representations, limiting realistic compositing, object reuse, and consistent manipulation. This limitation is particularly evident in video object insertion and video layer decomposition, where existing methods lack direct supervision for foreground layers that capture both objects and their associated visual effects. We introduce TriLayer, a triplet video dataset containing aligned composite--background--foreground videos, where the foreground layers include both object appearance and associated visual effects. With aligned triplet supervision, TriLayer enables explicit supervised learning of layered video representations for the first time. Building on this dataset, we propose DBL-Diffusion, a dual-branch diffusion framework that jointly models scene-level RGB content and RGBA foreground layers through cross-branch interaction during denoising. We instantiate the framework in two tasks: DBL-Insert for layered object insertion, which generates explicit RGBA layers for realistic compositing and flexible post-editing, and DBL-Decompose for video layer decomposition, which recovers foreground and background layers using triplet supervision. Experiments demonstrate that explicit layer modeling substantially improves both insertion fidelity and decomposition quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑