发表机构
Harbin Institute of Technology; Taobao, Alibaba Group(哈尔滨工业大学; 阿里巴巴集团淘宝)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ZoomDiff提出一种高保真扩散模型,通过增强条件引导和流对齐特征注入,解决双摄像头平滑变焦中的几何与色彩不连续问题,在合成和真实数据上超越现有方法。
AI 中文摘要
双摄像头之间的数字变焦过渡常常在几何结构和色彩一致性方面表现出明显的突变,从而降低用户体验。尽管最近的双摄像头平滑变焦(DCSZ)方法试图通过在DCSZ数据上微调帧插值(FI)模型来缓解这一问题,但它们难以应对大的跨视角差异和复杂的几何变换。考虑到扩散模型的生成先验适合解决这一问题,我们探索了其在DCSZ中的应用。然而,直接应用现有的基于扩散的FI模型仍会产生低保真度的过渡,原因是条件引导不足、VAE编码过程中的高频信息丢失以及时间一致性不足。为了解决这些问题,我们提出了ZoomDiff,一种高保真扩散模型,它在潜在空间和像素空间中利用双摄像头输入,以实现照片级逼真的过渡。具体来说,我们首先在多步去噪过程中加强双图像条件引导,以提高几何一致性。然后,我们将来自VAE编码器的流对齐多尺度特征注入VAE解码器,以恢复高频细节,并引入流引导的时间一致性监督,以产生更平滑的过渡。在合成和真实世界数据集上的大量实验表明,ZoomDiff在定量和定性方面均优于最先进的方法。项目页面:https://jiayi-hit.github.io/ZoomDiff.github.io/。
英文摘要
Digital zoom transitions between dual cameras often exhibit conspicuous discontinuities in geometric structure and chromatic consistency, degrading the user experience. While recent dual-camera smooth zoom (DCSZ) methods attempt to mitigate this by fine-tuning frame interpolation (FI) models on DCSZ data, they struggle with the large cross-view disparities and complex geometric transformations. Considering that the generative prior of diffusion models is suitable for addressing this problem, we explore their application to DCSZ. However, naively applying existing diffusion-based FI models still yields low-fidelity transitions due to insufficient conditional guidance, high-frequency information loss during VAE encoding, as well as inadequate temporal consistency. To address this, we propose ZoomDiff, a high-fidelity diffusion model that leverages dual-camera inputs in both latent and pixel spaces for photo-realistic transitions. Specifically, we first strengthen dual-image conditional guidance during the multi-step denoising process to improve geometric consistency. Then we inject flow-aligned multi-scale features from the VAE encoder into the VAE decoder to recover high-frequency details, where flow-guided temporal consistency supervision are introduced to produce more smooth transitions. Extensive experiments on both synthetic and real-world datasets demonstrate that ZoomDiff outperforms state-of-the-art methods quantitatively and qualitatively. Project page: https://jiayi-hit.github.io/ZoomDiff.github.io/.