FlexComposer:支持灵活轨迹控制的从图像到动态素材的统一视频合成框架
FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control
浏览论文内容
中文总结 AI 辅助
FlexComposer是支持灵活轨迹控制的统一视频合成框架,通过三项关键设计实现静态图像与动态素材的无缝整合,在视觉质量等指标上优于现有SOTA方法。
中文摘要 AI 辅助
生成式视频合成是将外部素材无缝插入现有视频序列的技术,对内容创作和视觉效果至关重要。但现有方法存在控制保真度的权衡问题:要么从静态图像生成幻觉运动,无法保留预动画素材的动态性;要么缺乏沿用户定义轨迹精确放置素材的细粒度空间控制。我们提出FlexComposer,这是一个将视频合成为轨迹引导条件生成任务的统一框架,可无缝整合静态图像和动态素材。该方法包含三项关键设计:(1)统一规范前景表示,将物体的内在运动与全局位移解耦,将异构输入标准化为稳定居中的潜在空间;(2)空间感知潜在注入策略,利用VAE潜在空间的平移等变性,通过无参数机制将规范特征传输到目标轨迹;(3)混合数据集与合成到真实课程,结合程序模拟、真实电影素材和生成数据,隐式学习物理上合理的光照与阴影协调。这种统一设计处理从产品照片到动态主体的多样化输入,实现高保真运动控制和环境整合,无需显式3D重建或辅助可学习适配器。大量实验表明,FlexComposer在视觉质量、时间一致性和轨迹贴合度上优于现有SOTA方法。
英文摘要
Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and visual effects. However, existing approaches suffer from a control-fidelity trade-off: they either hallucinate motion from static images, failing to preserve the dynamics of pre-animated assets, or lack fine-grained spatial control for precise asset placement along user-defined trajectories. We propose FlexComposer, a unified framework that standardizes video compositing as a trajectory-guided conditional generation task, enabling the seamless integration of both static images and dynamic footage. Our approach introduces three key designs: (1) a Unified Canonical Foreground Representation that decouples an object's intrinsic motion from its global displacement, standardizing heterogeneous inputs into a stabilized, centered latent space; (2) a Spatial-Aware Latent Injection strategy that exploits the translation equivariance of VAE latent spaces to transport canonical features onto target trajectories via a parameter-free mechanism; and (3) a Hybrid Dataset and Synthetic-to-Real Curriculum that synergizes procedural simulation, real-world cinematic footage, and generative data to implicitly learn physically plausible illumination and shadow harmonization. This unified design handles diverse inputs from product photos to dynamic subjects achieving high-fidelity motion control and environmental integration without the need for explicit 3D reconstruction or auxiliary learnable adapters. Extensive experiments demonstrate that FlexComposer outperforms state-of-the-art methods in visual quality, temporal consistency, and trajectory adherence.
发表机构
- HKUST(香港科技大学)
- ZJU(浙江大学)
- CUHK(香港中文大学)
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。