保持图像外扩中主体清晰度的多尺度小波监督
Preserving Subject-Clarity in Image Outpainting with Multiscale Wavelet Supervision
浏览论文内容
中文总结 AI 辅助
针对图像外扩中主体清晰度下降问题,提出结合VLM语义条件与多尺度小波监督的框架,并构建主体中心数据流程,在四个基准上显著降低DreamSim误差和FID。
中文摘要 AI 辅助
商业和广告图像经常受到构图不佳、主体部分被裁剪、文本或标志被截断以及上下文不足的影响,这些都会降低主体清晰度,即图像清晰传达其主要主体的能力。图像外扩通过扩展图像边界并恢复缺失的内容和上下文,提供了一种可扩展的解决方案。然而,现有的基于扩散的外扩方法往往产生视觉上合理的补全,但通过结构不一致、语义漂移或细粒度细节的丢失而降低了主体保真度。为了解决这一限制,我们提出了一种主体清晰度外扩框架,该框架将视觉语言模型(VLM)引导的语义条件与多尺度小波监督相结合,用于主体局部细节的保留。为了支持训练,我们开发了一个以主体为中心的数据整理流程,从广告和自然图像中构建主体相交的外扩对。所得到的目标函数不引入额外的推理成本,并设计为与基于扩散的主干网络兼容。在四个广告和自然图像基准上,我们的方法提高了主体清晰度,与匹配的监督微调相比,平均将主体中心的DreamSim误差和FID分别降低了3.0%和2.4%,与每个数据集上最强的最先进方法相比,分别降低了10.8%和7.7%。
英文摘要
Commercial and advertising images are frequently affected by poor framing, partially cropped subjects, truncated text or logos, and insufficient context, all of which can reduce subject clarity, i.e., the ability of an image to clearly communicate its primary subject. Image outpainting offers a scalable solution by extending image boundaries and recovering missing content and context. However, existing diffusion-based outpainting methods often produce visually plausible completions while degrading subject fidelity through structural inconsistencies, semantic drift, or loss of fine-grained detail. To address this limitation, we propose a subject clarity outpainting framework that combines vision-language model (VLM)-guided semantic conditioning with multiscale wavelet supervision for subject-localized detail preservation. To support training, we develop a subject-centric data curation pipeline that constructs subject-intersecting outpainting pairs from advertising and natural images. The resulting objective introduces no additional inference cost and is designed to be compatible with diffusion-based backbones. Across four advertising and natural-image benchmarks, our method improves subject clarity, reducing subject-centered DreamSim error and FID on average by 3.0% and 2.4% over matched supervised fine-tuning, and by 10.8% and 7.7% over the strongest state-of-the-art approach per dataset, respectively.
发表机构
- Microsoft(微软)
- Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。