arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07463cs.CVcs.LG

MirrorWorld:利用视频扩散模型生成镜像反射

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

Youjun Zhao, Alex Warren, Gary K. L. Tam, Rynson W. H. Lau

首次发表
浏览论文内容

中文总结 AI 辅助

MirrorWorld是一种反射感知视频修复框架,通过语义关系蒸馏和几何变换对齐建模场景与镜像的关系,在视频镜像反射重建任务上优于现有方法。

中文摘要 AI 辅助

近期视频扩散模型(VDMs)的进展已实现高保真视频合成,但生成镜像反射仍具挑战性,因为镜像内的内容必须与周围场景保持一致。现有VDMs并非专门用于建模场景与镜像的关系,这会导致反射内容不正确或空间排列不一致。我们观察到镜像反射生成涉及两个互补挑战:确定应反射的场景内容,以及反射内容在镜像区域内的空间排列方式。基于此观察,我们提出MirrorWorld,这是一个反射感知的视频修复框架,在生成过程中建模场景与镜像的关系。具体而言,我们引入语义关系蒸馏(SRD),从冻结的视觉基础模型中传递关系信息,以鼓励可见场景内容与镜像区域之间的语义关联。我们进一步提出几何变换对齐(GTA),学习一种变换来指导反射内容的空间排列。这两个组件发挥互补作用:SRD建模应反射的内容,GTA建模其排列方式。为促进该问题的研究,我们通过将四个现有视频镜像数据集重新用于统一的反射重建任务,构建了一个视频镜像反射生成基准。实验结果表明,MirrorWorld比代表性的基于图像的反射生成方法和强大的视频修复基线实现了更好的反射重建质量。

英文摘要

Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VDMs are not specifically designed to model scene-to-mirror relationships, which can lead to reflections with incorrect content or inconsistent spatial arrangements. We observe that mirror reflection generation involves two complementary challenges: determining what scene content should be reflected and how the reflected content should be spatially arranged within the mirror region. Motivated by this observation, we propose MirrorWorld, a reflection-aware video inpainting framework that models scene-to-mirror relationships during generation. Specifically, we introduce Semantic Relation Distillation (SRD), which transfers relational information from a frozen visual foundation model to encourage semantic associations between visible scene content and mirror regions. We further propose Geometric Transformation Alignment (GTA), which learns a transformation that guides the spatial arrangement of reflected content. The two components play complementary roles, with SRD modeling what should be reflected and GTA modeling how it should be arranged. To facilitate research on this problem, we construct a benchmark for video mirror reflection generation by repurposing four existing video mirror datasets into a unified reflection reconstruction task. Experimental results show that MirrorWorld achieves improved reflection reconstruction quality over representative image-based reflection generation methods and strong video inpainting baselines.

补充信息

↑