发表机构
Durham University(杜伦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建MDTD-Art基准及MDTD-Art数据集,对比各类模型,发现图像编辑模型在艺术图像修复中优于专用修复架构,结构化提示可放大其性能优势。
AI 中文摘要
严重退化视觉媒体的修复仍是一项艰巨挑战,现有方法常生成不自然的纹理与内容,难以保留色彩和纹理,或无法利用部分保留的图像信息。现有修复基准假设退化算子已知,无法捕捉艺术损伤如裂纹、污渍及色彩/纹理偏差等复杂特征。我们引入了针对语义半透明图像媒体退化的盲修复可控基准,附带新的公开退化alpha纹理掩码数据集MDTD-Art。我们提出了新的数据集与基准,在不同掩码不透明度水平下评估了最先进的通用修复模型、图像编辑模型及视觉语言模型。实验表明,对于任意退化,图像编辑模型始终优于专用修复架构,通过强调细节保留和结构一致性的结构化提示工程,性能提升被放大。这些发现表明,可恢复的语义信息和提示可控性是艺术图像修复的关键因素。
英文摘要
Restoring severely degraded visual media still remains a formidable challenge, as existing methods often hallucinate unnatural textures and contents, struggle with preserving color and texture, or fail to leverage partially retained image information. Existing restoration benchmarks assume known degradation operators and fail to capture the complex characteristics of artistic damage such as cracks, stains, and color/texture deviation. We introduce a controlled benchmark for blind restoration of semantic, semi-transparent image media degradations, accompanied by a new, publicly open degradation alpha texture mask dataset MDTD-Art. We present a new dataset and benchmark evaluating state-of-the-art universal restoration models against image editing and vision-language models across varying mask opacity levels. Our experiments demonstrate that image editing models consistently outperform specialized restoration architectures for arbitrary degradations, with performance gains amplified by structured prompt engineering emphasizing detail preservation and structural consistency. These findings position recoverable semantic information and prompt controllability as critical factors in art image restoration.
Comments15 pages, 6 figures, 6 tables. Preprint submitted to Elsevier Journal of Visual Communication and Image Representation