arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25716cs.CV

FoMo:生成轨迹中的分叉时刻作为感知距离

FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance

Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出FoMo,利用扩散模型生成轨迹中的分叉时刻作为感知距离标签,无需人工标注即可训练基于参考的图像质量评估指标,在多个基准上超越人工标注数据集。

中文摘要 AI 辅助

基于参考的图像质量评估(IQA)指标旨在反映人类如何感知一对图像之间的感知距离。为了学习人类视觉系统(HVS)的运作方式,最近的基于参考的IQA指标严重依赖人工标注数据。基于平均意见分数(MOS)的逐点评分,即为每幅图像分配一个标量质量值,虽然有利于标注,但在大规模收集时成本过高,并且由于人类判断的不一致性而众所周知地存在噪声。作为替代方案,二选一强制选择(2AFC)成对标签因其可靠性和效率而受到青睐,但它们仅捕获成对之间的相对比较。在本文中,我们提出了一种全自动数据生成管道,无需任何人工标注即可生成图像对之间的逐点感知距离标签。我们的方法利用扩散模型的生成动力学作为感知距离的代理,其中图像的粗略结构在早期时间步生成,而精细细节在后期时间步生成。在生成过程中早期分叉的图像仅共享粗略结构,在感知上相距较远;晚期分叉的图像仅在精细细节上有所不同。我们证明了扩散轨迹与人类视觉系统高度一致,并使用这一分叉时刻FoMo作为参考基础的距禒标签来监督基于参考的IQA指标的训练。支持任意图像对之间通用比较的逐点标签,实现了信息丰富的训练目标。跨多种骨干架构的大量实验证实了我们生成管道的有效性,在多个基准上优于人工标注数据集。

英文摘要

Reference-based image quality assessment (IQA) metrics aim to reflect how humans perceive the perceptual distance between a pair of images. To learn how the human visual system (HVS) operates, recent reference-based IQA metrics heavily rely on human-annotated data. Mean opinion score (MOS)-based pointwise scoring, which assigns a scalar quality value per image, is preferable for annotation but is prohibitively expensive to collect at scale and is known to be noisy due to inconsistent human judgments. As an alternative, two-alternative forced choice (2AFC) pairwise labels have gained popularity due to their reliability and efficiency, but they capture only relative comparisons between pairs. In this paper, we propose a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation. Our approach exploits the generative dynamics of diffusion models as a perceptual distance proxy, where the coarse structure of an image is generated in the early timesteps and the fine details are generated in the later timesteps. Images that fork early in the generation process share only coarse structure and are perceptually far apart; images that fork late differ only in fine detail. We demonstrate that the diffusion trajectory aligns well with the human visual system, and use this forking moment, FoMo, as a reference-grounded distance label to supervise the training of a reference-based IQA metric. The pointwise labels, which support universal comparison between arbitrary image pairs, enable an information-rich training objective. Extensive experiments across diverse backbone architectures confirm the effectiveness of our generation pipeline, outperforming human-annotated datasets in multiple benchmarks.

发表机构

  • Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑