基于潜在感知自适应掩码的源无关图像翻译
Source-Agnostic Image Translation Based on Latent Aware Adaptive Masking
浏览论文内容
中文总结 AI 辅助
本研究提出一种源无关图像翻译框架,通过潜在感知自适应掩码方案,在AFHQ、Celeba-HQ数据集上优于现有无监督图像到图像方法,无需专门训练即可实现跨源分布的无缝翻译。
中文摘要 AI 辅助
在本研究中,我们提出一种源无关框架,该框架通过计算预训练扩散模型在每个潜在时间步的预测差异,在反向扩散过程中动态优化二进制掩码。与依赖固定阈值不同,我们的方法引入一种时间依赖的统计阈值方案,该方案源自目标分布的潜在噪声图像中预测差异的经验均值和标准差。这使得掩码能够适应模型在不同噪声水平下变化的预测置信度,有效隔离特定域区域,同时保留全局结构一致性。在AFHQ和Celeba-HQ数据集上的实验结果表明,我们的方法在真实性(FID、KID)和忠实性(SSIM、LPIPS)方面均优于最先进的无监督图像到图像方法。由于仅需目标域的预训练模型,我们的方法无需任何专门训练即可实现精确、自动的定位,并在不同源分布间实现无缝翻译。项目源代码可访问:this https URL
英文摘要
In this work, we propose a source-agnostic framework that dynamically refines a binary mask throughout the reverse diffusion process by computing the discrepancies of a pretrained diffusion model's prediction for each latent time step. Rather than relying on a fixed threshold, our method introduces a time-dependent statistical thresholding scheme derived from the empirical mean and standard deviation of prediction discrepancies across the latent noisy images from the target distribution. This allows the mask to adapt to the model's varying predictive confidence at different noise levels, effectively isolating domain-specific regions while preserving global structural coherence. Experimental results on the AFHQ and Celeba-HQ datasets demonstrate that our approach outperforms state-of-the-art unsupervised Image-to-Image methods in both realism (FID, KID) and faithfulness (SSIM, LPIPS). By requiring only a pretrained model of the target domain, our approach enables precise, automated localization and seamless translation across diverse source distributions without any specialized training. The project source code is available at: https://github.com/dtoma95/PM-Edit
发表机构
- Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。