arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向对视觉皮层和扩散模型中推理的机制性理解

Toward a mechanistic understanding of inference in visual cortex and diffusion models

Zeyu Yun, Alexander Belsten, Dasheng Bi, Zahra Kadkhodaie, Yubei Chen, Bruno A. Olshausen

arXiv 2607.15693首次发表:更新:

发表机构

Dept. of Electrical Engineering and Computer Science, UC Berkeley; Helen Wills Neuroscience Institute, UC Berkeley; Herbert Wertheim School of Optometry and Vision Science, UC Berkeley; Redwood Center for Theoretical Neuroscience, UC Berkeley; Flatiron Institute; Dept. of Electrical and Computer Engineering, UC Davis(加州大学伯克利分校电气工程与计算机科学系; 加州大学伯克利分校海伦·威尔斯神经科学研究所; 加州大学伯克利分校赫伯特·韦特海姆视觉科学学院; 加州大学伯克利分校红木理论神经科学中心; Flatiron研究所; 加州大学戴维斯分校电气与计算机工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在理解视觉皮层和扩散模型中推理,提出基于稀疏编码的模型,用去噪分数匹配等训练,其性能出色,能揭示递归动力学机制,还能断开部分潜在变量与视觉输入连接,连接神经科学与机器学习领域。

AI 中文摘要

我们描述了一种初级视觉皮层(V1)中的感知推理模型,它等同于一个最小扩散模型,其功能可从参数中轻易理解。该模型基于稀疏编码,对潜在变量采用非因子先验,形式为无约束的成对交互矩阵,将标准稀疏编码推理扩展到一般递归动力系统。我们使用去噪分数匹配目标和隐式微分有效训练这些递归动力学。在自然图像上训练后,学习到的交互矩阵反映了V1浅层水平连接的结构。此模型具有出色的去噪性能,在极端视觉模糊中恢复图像特征,在泛化方面接近标准黑箱扩散架构。因其简单性,网络雅可比矩阵可直接根据潜在变量间的交互矩阵分解,揭示递归动力学如何在连续自然结构变形族上赋予高概率。有趣的是,在这个回路中,很大一部分潜在变量完全与视觉输入断开连接,形成层次表示,增强图像特征的全局一致性。该模型和结果连接了两个不同领域:为神经科学生成关于感知推理任务中递归神经回路功能连接的具体可测试假设;为机器学习阐明扩散模型学习的内部机制,使其能从有限训练集中生成无限多新图像。

英文摘要

We describe a model of perceptual inference in primary visual cortex (V1) equivalent to a minimal diffusion model whose function can be readily understood from its parameters. The model is based on sparse coding with a non-factorial prior over latent variables in the form of an unconstrained, pairwise interaction matrix, extending standard sparse coding inference to a general recurrent dynamical system. We efficiently train these recurrent dynamics using a denoising score-matching objective and implicit differentiation. After training on natural images, the learned interaction matrix mirrors the structure of horizontal connections in superficial layers of V1 that link neurons of similar orientation tuning. This model exhibits exceptionally good denoising performance, restoring image features such as extended contours amid extreme visual ambiguity, nearly matching the behavior of standard, black-box diffusion architectures in generalization regime. Owing to the model's simplicity, the network's Jacobian can be decomposed directly in terms of the interaction matrix between latent variables, revealing mechanistically how the recurrent dynamics assign high probability over a continuous family of natural structural deformations. Intriguingly, within this circuit, a large fraction of latent variables learn to disconnect from visual input altogether, essentially forming a hierarchical representation that appears to enforce global consistency among image features. Together, the model and results bridge two distinct domains: for neuroscience, it generates concrete, testable hypotheses regarding functional connectivity in recurrent neural circuits during perceptual inference tasks; for machine learning, it elucidates the internal mechanisms learned by diffusion models that allow them to generate infinitely many novel images from a finite training set.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑