arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

去噪模型在不同架构中开发出类似人类的感知错觉表示

Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures

Gautam Ranka, Paras Chopra

arXiv 2607.17138首次发表:更新:

发表机构

Lossfunk(Lossfunk)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨自然图像训练的深度神经网络在亮度错觉输出与人类一致,发现去噪模型在特定内部层有错觉敏感表示,确定相关层和通道,表明去噪目标更关键,虽激活跟踪人类感知模型,但注入表示无像素偏移,此为去噪视觉模型感知表示的首次此类特征描述。

AI 中文摘要

在自然图像上训练的深度神经网络在亮度错觉方面产生了与人类观察者一致的输出。虽然这种现象在各种架构中都有记录,但目前所有证据都是在输出层面衡量的。我们表明,去噪模型在不同架构的特定内部层开发出错觉敏感表示,确定了区分错觉与物理匹配控制区域的层和通道。去噪目标比架构更重要。这些激活跟踪人类亮度感知的心理物理模型,通过通道消融提供因果证据,表明错觉敏感通道影响内部信号,但注入这些表示不会产生可测量的像素偏移,我们称这种表示为感知幻影。

英文摘要

Deep neural networks trained on natural images are shown to produce outputs consistent with human observers for brightness illusions. While this phenomenon has been documented across architectures, all evidence, to date, is measured at the output level: restored pixels, decoded trajectories, or classification decisions. Whether these models actually represent illusions internally, and if so where and how, remains unknown. We show that denoising models develop illusion-sensitive representations at specific internal layers, across varied architectures. Specifically, we identify the layers and channels that discriminate illusory from physically matched control regions. We show that the denoising objective is a more important driver of the effect than the architecture. On domain-appropriate stimuli, these activations track a validated psychophysical model of human brightness perception (FLODOG; Spearman $ρ\geq 0.70$) and scale monotonically with parametric illusion strength. Leveraging these findings, we provide causal evidence via channel ablation showing that illusion-sensitive channels specifically and substantially affect the internal signal. Yet injecting these representations into the generation pipeline produces no measurable pixel shift across all tested architectures; we term such representations perceptual phantoms: active in internal processing yet invisible to any output-based evaluation. While related internal-output dissociations have been characterized in language models, this is the first such characterization for perceptual representations in denoising vision models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑