arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36391q-bio.NCcs.LG

感受野约束下的人类早期和中间视觉皮层刺激优化

Receptive-field-constrained stimulus optimization for human early and intermediate visual cortex

  • Carnegie Mellon University(卡内基梅隆大学)
  • University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

Junru Zhao, Hanfei Guo, Andrew Luo, Margaret M. Henderson

AI总结:

针对早期和中间视觉皮层小感受野的挑战,提出RF-DiVE和RF-GO两种受pRF约束的MEI生成框架,在V1-hV4区域生成结构一致的MEI,其预测响应优于自然图像,实现数据驱动的特征选择性表征。

AI中文摘要:

感觉神经科学中的一个持续挑战是表征皮层群体所编码的特征维度。近期方法通过为目标神经群体合成最激发输入(MEI),以数据驱动的方式探测特征选择性。虽然该方法已成功应用于使用fMRI数据的人类高级视觉皮层,但由于感受野尺寸较小,为早期和中期视网膜拓扑视觉区域生成MEI需要额外的建模约束。为应对这一挑战,我们引入了两种新颖的MEI生成框架:感受野扩散视觉探索(RF-DiVE)和感受野梯度优化(RF-GO)。两种方法均使用群体感受野(pRF)约束的体素级编码模型;RF-DiVE将其与预训练的潜在扩散模型相结合,而RF-GO则使用正则化梯度上升。当应用于视网膜拓扑定义的V1-hV4区域中的单个体素时,使用自然场景数据集的数据,我们获得的MEI在pRF内展现出一致的结构,表明对轮廓、颜色和纹理等局部特征的选择性。我们系统比较了使用两种编码骨干的RF-DiVE和RF-GO生成的MEI,并使用独立编码模型对MEI的预测响应进行了硅内验证。在所有方法和所有视觉区域中,MEI引发的模型预测响应高于最具激活性的自然图像。我们进一步发现,生成框架和编码骨干的选择对MEI特性产生不同影响,包括其视觉外观、结构可解释性和跨模型泛化性。这些结果为跨人类视觉皮层进行空间和特征选择性的数据驱动表征提供了一种新方法。

英文摘要:

An ongoing challenge in sensory neuroscience is to characterize the feature dimensions encoded by cortical populations. Recent approaches probe feature selectivity in a data-driven way, by synthesizing a most-exciting-input (MEI) for a target neural population. While this approach has been successfully applied to human higher visual cortex using fMRI data, generating MEIs for early- and mid-level retinotopic visual areas requires additional modeling constraints due to small receptive field sizes. To address this challenge, we introduce two novel MEI generation frameworks, Receptive Field Diffusion for Visual Exploration (RF-DiVE) and Receptive Field Gradient Optimization (RF-GO). Both methods use a population receptive field (pRF)-constrained voxelwise encoding model; RF-DiVE combines this with a pretrained latent diffusion model, while RF-GO uses regularized gradient ascent. When applied to single voxels in retinotopically defined areas V1-hV4, using data from the Natural Scenes Dataset, we obtain MEIs that exhibit consistent structure within the pRF, suggesting selectivity for local features like contour, color, and texture. We systematically compare MEIs generated by RF-DiVE and RF-GO using two encoding backbones, performing in-silico validation of predicted responses to MEIs using independent encoding models. Across all methods and all visual areas, MEIs elicit higher model-predicted responses than the most activating natural images. We further find that the choice of generation framework and encoding backbone differentially affects MEI properties, including their visual appearance, structural interpretability, and cross-model generalizability. These results offer a new approach for performing data-driven characterization of spatial and feature selectivity across human visual cortex.

补充信息

↑