arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

域自适应零样本图像增强:基于局部约束扩散引导

Domain-adaptive Zero-Shot Image Enhancement via Locality-Constrained Diffusion Guidance

Theresa Neubauer, Dimitrios Lenis, Astrid Berg, Maria Wimmer, Gaia Romana De Paolis, Philip Matthias Winter, David Major, Johannes Novotny, Ariharasudhan Muthusami, Katja Bühler

arXiv 2609.35289首次发表:更新:

发表机构

VRVis GmbH(VRVis 有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出LocDiff,一种局部约束扩散引导方法,在零样本设置下实现跨域图像增强,平衡目标域真实感与源域关键特征保留,并在艺术转照片和低质量超声增强任务上验证了其优越性。

AI 中文摘要

去噪扩散概率模型在无条件图像生成方面展现了卓越的性能。为了生成具有所需语义的图像,近期研究通过在扩散采样过程中施加引导约束来限制解空间。然而,对于跨不同域的图像增强,这些方法难以平衡两个主要需求:在目标域中看起来逼真(照片级真实图像)以及保留源域的相关特征,例如低质量渲染或艺术绘画。在此,小的局部变化可能完全改变图像的保真度,而其他区域的大变化可能无关紧要。我们引入了LocDiff,一种用于图像增强的局部约束引导方法,作为预训练扩散模型的零样本扩展,确保在域适应过程中保留关键特征。通过这种方式,我们保留了重要的局部特征,同时允许较不关键的区域保持不受约束,且不干扰相关区域的引导过程。我们在两个不同的域偏移任务上评估了我们的方法:对于艺术到照片的转换,我们在完全零样本设置下应用该方法,在生成照片级真实细节的同时保留绘画中的面部身份。对于增强低质量胎儿超声渲染,我们展示了带有辅助先验对齐的零样本推理。这里的目标是人为添加高分辨率特征并生成照片级真实的超声渲染,这是一个不存在真实分布的目标域。我们的实验结果表明,与最先进的方法相比,LocDiff实现了有利的真实感-保真度权衡,从而实现可控的跨域增强。

英文摘要

Denoising Diffusion Probabilistic Models have shown remarkable performance in unconditional image generation. In order to generate images with desired semantics, recent works have restricted the solution space by using guidance constraints in the diffusion sampling process. However, for image enhancement across different domains, these methods struggle to balance two main requirements: looking realistic in the target domain (photorealistic images) and preserving relevant features of the source domain, e.g., low-quality renderings or art paintings. Here, small local changes can alter the fidelity of the image completely, while large changes in other regions might be insignificant. We introduce LocDiff, a locality-constrained guidance method for image enhancement, which serves as a zero-shot extension to pre-trained diffusion models, ensuring the preservation of critical features during domain adaptation. In this way, we retain important local features, while allowing less critical regions to remain unconstrained and not interfere with the guidance process for relevant regions. We evaluate our method on two different domain-shift tasks: For art-to-photo translation, we apply the method in a fully zero-shot setting, preserving facial identity from paintings while generating photorealistic details. For enhancing low-quality fetal ultrasound renderings, we demonstrate zero-shot inference with auxiliary prior alignment. Here, the objective is to artificially add high-resolution characteristics and produce photorealistic ultrasound renderings, a target domain for which no ground truth distribution exists. Our experimental results demonstrate that LocDiff achieves favorable realism-faithfulness trade-offs compared to state-of-the-art methods, enabling controllable cross-domain enhancement.

CommentsAccepted manuscript. The final version is published in Computers & Graphics

Journal refComputers & Graphics, Volume 137, 2026, 104607

DOI:10.1016/j.cag.2026.104607

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑