arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29627cs.CV

FlexComposer:支持灵活轨迹控制的从图像到动态素材的统一视频合成框架

FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

Songchun Zhang, Sitong Guo, Xianghao Kong, Pengwei Liu, Yuwei Guo, Lvmin Zhang, Anyi Rao

首次发表
浏览论文内容

中文总结 AI 辅助

FlexComposer是支持灵活轨迹控制的统一视频合成框架,通过三项关键设计实现静态图像与动态素材的无缝整合,在视觉质量等指标上优于现有SOTA方法。

中文摘要 AI 辅助

生成式视频合成是将外部素材无缝插入现有视频序列的技术,对内容创作和视觉效果至关重要。但现有方法存在控制保真度的权衡问题:要么从静态图像生成幻觉运动,无法保留预动画素材的动态性;要么缺乏沿用户定义轨迹精确放置素材的细粒度空间控制。我们提出FlexComposer,这是一个将视频合成为轨迹引导条件生成任务的统一框架,可无缝整合静态图像和动态素材。该方法包含三项关键设计:(1)统一规范前景表示,将物体的内在运动与全局位移解耦,将异构输入标准化为稳定居中的潜在空间;(2)空间感知潜在注入策略,利用VAE潜在空间的平移等变性,通过无参数机制将规范特征传输到目标轨迹;(3)混合数据集与合成到真实课程,结合程序模拟、真实电影素材和生成数据,隐式学习物理上合理的光照与阴影协调。这种统一设计处理从产品照片到动态主体的多样化输入,实现高保真运动控制和环境整合,无需显式3D重建或辅助可学习适配器。大量实验表明,FlexComposer在视觉质量、时间一致性和轨迹贴合度上优于现有SOTA方法。

英文摘要

Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and visual effects. However, existing approaches suffer from a control-fidelity trade-off: they either hallucinate motion from static images, failing to preserve the dynamics of pre-animated assets, or lack fine-grained spatial control for precise asset placement along user-defined trajectories. We propose FlexComposer, a unified framework that standardizes video compositing as a trajectory-guided conditional generation task, enabling the seamless integration of both static images and dynamic footage. Our approach introduces three key designs: (1) a Unified Canonical Foreground Representation that decouples an object's intrinsic motion from its global displacement, standardizing heterogeneous inputs into a stabilized, centered latent space; (2) a Spatial-Aware Latent Injection strategy that exploits the translation equivariance of VAE latent spaces to transport canonical features onto target trajectories via a parameter-free mechanism; and (3) a Hybrid Dataset and Synthetic-to-Real Curriculum that synergizes procedural simulation, real-world cinematic footage, and generative data to implicitly learn physically plausible illumination and shadow harmonization. This unified design handles diverse inputs from product photos to dynamic subjects achieving high-fidelity motion control and environmental integration without the need for explicit 3D reconstruction or auxiliary learnable adapters. Extensive experiments demonstrate that FlexComposer outperforms state-of-the-art methods in visual quality, temporal consistency, and trajectory adherence.

发表机构

  • HKUST(香港科技大学)
  • ZJU(浙江大学)
  • CUHK(香港中文大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑