发表机构
National Cheng Kung University; Harvard Medical School(成功大学; 哈佛医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对有损压缩图像复原中保真度、人类感知与机器偏好的权衡问题,提出两阶段架构FDIR,实验证实其在三者间取得更优平衡。
AI 中文摘要
图像复原质量可从三个互补维度评估:像素级保真度、人类感知及下游机器偏好。然而,现有有损压缩复原方法最多仅优化其中一项标准:面向保真度的模型常向条件均值回归,生成过平滑输出;而生成式方法会生成看似合理但与事实不符的纹理,既降低与真值的保真度,又损害下游任务精度。为应对这三者的权衡,我们提出FDIR,这是一种两阶段架构,通过互补归纳偏置化解冲突需求:质量引导单步流匹配(QO-Flow)通过单次前向传播在隐空间恢复全局语义结构;流条件细节精修(FCDR)在像素空间确定性地恢复高频纹理并抑制生成式幻觉。大量实验表明,FDIR实现了更优的保真度,具备良好的感知-保真度平衡及有竞争力的机器偏好。
英文摘要
Image restoration quality can be evaluated along three complementary facets: pixel-level fidelity, human perception, and downstream machine preference. However, existing lossy compression restoration methods optimize for at most one of these criteria: fidelity-oriented models often regress toward conditional means and produce over-smoothed outputs, while generative approaches hallucinate plausible but factually incorrect textures that degrade both ground-truth fidelity and downstream task accuracy. To navigate this three-way tradeoff, we propose FDIR, a two-stage architecture that decouples the conflicting demands through complementary inductive biases: Quality-Guided One-Step Flow Matching (QO-Flow) recovers global semantic structure in latent space via a single forward pass, while Flow-Conditioned Detail Refinement (FCDR) deterministically restores high-frequency textures and suppresses generative hallucinations in pixel space. Extensive experiments demonstrate that FDIR achieves superior fidelity, with a favorable perceptual-fidelity balance and competitive machine preference.
Comments19 pages, 8 figures, 13 tables