发表机构
School of Data Science, The Chinese University of Hong Kong, Shenzhen; School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院; 香港中文大学(深圳)理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Step-PI方法,借助冻结文本到图像扩散模型,无需训练即可实现跨域图像修复,在多个数据集上优于同类无训练基线,为跨域图像修复提供了补充方案。
AI 中文摘要
我们证明,冻结的通用文本到图像扩散模型可在三种评估的自然图像域中执行条件图像修复,且仅需固定的控制器配置,无需图像修复特定的权重训练、数据集特定的权重适配或学习到的图像修复特定条件通道。Step-PI将已知区域投影扩展为包含边界-内部潜在反馈、持久PI状态,以及沿反向轨迹调节控制器信号的预定义四场释放调度。控制器仅在Main35不相交的CelebA-HQ试点数据上开发,可原封不动地迁移到AFHQ和Places2。在同一3500个案例的两次场匹配比较中,添加持久状态并将均匀释放替换为预定义调度,分别提升了全部15个数据集-指标单元;两次比较中所有五个指标的95%自助法区间均排除零值。在描述性原生路径比较中,Step-PI在所有五个等数据集宏观指标上优于LanPaint和PILOT(使用普通SD1.5的最接近的评估无训练基线)。经图像修复训练的系统保留了绝对指标优势,但依赖大量图像修复特定的离线优化。我们的方法为通过测试时潜在控制将冻结通用文本到图像模型重新用于跨域图像修复提供了一种补充途径。
英文摘要
We show that a frozen generic text-to-image diffusion model can perform conditional inpainting across three evaluated natural-image domains with one fixed controller configuration, without inpainting-specific weight training, dataset-specific weight adaptation, or learned inpainting-specific conditioning channels. Step-PI augments known-region projection with boundary-interior latent feedback, persistent PI state, and a predefined four-field release schedule that modulates controller signals along the reverse trajectory. Developed only on Main35-disjoint CelebA-HQ pilots, the controller transfers unchanged to AFHQ and Places2. Across two field-identical comparisons on the same 3,500 cases, adding persistent state and replacing uniform release with the predefined schedule each improve all 15 dataset-metric cells; 95% bootstrap intervals exclude zero for all five metrics in both comparisons. In descriptive native-route comparisons, Step-PI leads LanPaint and PILOT (the closest evaluated training-free baselines using vanilla SD1.5) on all five equal-dataset macro metrics. Inpainting-trained systems retain the absolute metric leads but rely on substantial inpainting-specific offline optimization. Our method provides a complementary approach for repurposing a frozen generic text-to-image model for cross-domain inpainting through test-time latent control.
Comments30 pages, 11 figures, including supplementary material