arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成先验是否与人类自然感感知对齐?

Do Generative Priors Align with Human Naturalness Perception?

Taiki Fukiage

arXiv 2610.09928首次发表:更新:

发表机构

Communication Science Laboratories, NTT, Inc.(NTT通信科学实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过25个生成模型的预测损失差异,发现生成先验能捕捉人类对自然感的选择性敏感性与耐受性,并揭示其与通用违规检测的解耦关系。

AI 中文摘要

视觉生成模型被训练用于捕捉自然图像的概率分布,然而其原生先验是否反映了支配人类对图像自然感感知的规律仍是一个悬而未决的问题。在此,我们通过25个开源图像和视频生成器的原生预测误差来探究这些先验。由于原始单图像损失受场景内容和视觉复杂性的主导,我们使用内容保持的配对关系干预来评估方向性损失差异,这些干预选择性地破坏面部构型或物理光照一致性,同时限制低级图像统计的变化。在这两个领域中,这些损失差异再现了类似人类的选择性敏感性和耐受性,捕捉了经典的面孔撒切尔效应和对光照不一致性的形状依赖性反应。值得注意的是,这些损失差异可靠地追踪了个体刺激对之间人类自然感判断的连续梯度(在面孔上峰值相关系数r = .84,在物理场景上为.64),并且在控制了冻结视觉编码器的特征距离和标准图像质量指标后,仍保留了独特的人类对齐信号。我们还发现,虽然对这些违规行为的整体敏感性与跨模型的人类对齐广泛共变,但两者在去噪调度过程中系统地解耦,对齐峰值早于敏感性,揭示出类似人类的自然感判断与通用违规检测分离。综合这些发现表明,学习视觉分布产生的生成损失景观捕捉了人类自然感感知的不同方面。

英文摘要

Visual generative models are trained to capture the probability distributions of natural images, yet whether their native priors reflect the regularities governing human perception of image naturalness remains an open question. Here, we probe these priors through native prediction errors across 25 open image and video generators. Because raw single-image losses are dominated by scene content and visual complexity, we evaluate directional loss differences using content-preserving, paired relational interventions that selectively disrupt facial configurations or physical illumination consistency while limiting changes in low-level image statistics. Across both domains, these loss differences reproduce human-like selective sensitivities and tolerances, capturing the classic Thatcher effect on faces and shape-dependent responses to illumination inconsistencies. Notably, these loss differences reliably track continuous gradations of human naturalness judgments across individual stimulus pairs (peaking at $r = .84$ on faces and $.64$ on physical scenes) and retain unique human-aligned signals even after controlling for feature distances from frozen vision encoders and standard image quality metrics. We also find that while overall sensitivity to these violations broadly covaries with human alignment across models, the two systematically decouple along denoising schedules, with alignment peaking earlier than sensitivity, revealing that human-like naturalness judgments dissociate from generic violation detection. Together, these findings demonstrate that learning visual distributions yields generative loss landscapes that capture distinct aspects of human naturalness perception.

Comments62 pages, 35 figures, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑