发表机构
EPFL; Johns Hopkins University(苏黎世联邦理工学院; 约翰霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出神经引导视频合成框架NEvo,通过进化搜索优化视频刺激以激活目标脑区,揭示视觉皮层动态选择性,并发现侧流对社会动态特征的敏感性。
AI 中文摘要
人脑通过层级组织、功能特化的区域处理动态视觉输入。尽管最近的脑编码模型可以合成最优刺激来探测不同脑区的选择性,但先前工作主要局限于静态图像,动态视觉处理尚未充分探索。我们提出了一种新颖的神经引导视频合成框架,该框架生成针对视觉皮层目标脑区优化的刺激。我们的方法在结构化提示空间中进行进化搜索,由预测视频输入体素级响应的动态编码模型引导。通过最大化目标ROI的预测活动,该框架高效发现超激活动态刺激,这些刺激始终优于手工制作的定位器视频。合成的视频恢复了腹侧、背侧和侧通路的已知选择性,并进一步揭示了时间动态敏感性的系统性差异。一项搜索光分析提供了对沿侧流逐渐复杂的社会动态特征进展的新见解,并通过合成抽象、非自然刺激的探测进一步支持。综合来看,我们的框架能够对动态视觉选择性进行计算机模拟探索,并为体内实验提供新预测。
英文摘要
The human brain processes dynamic visual input through hierarchically organized, functionally specialized regions. While recent in silico brain encoding models can synthesize optimal stimuli to probe selectivity in different brain regions, prior work has been largely limited to static images, leaving dynamic visual processing underexplored. We introduce a novel neural-guided video synthesis framework that generates stimuli optimized for target brain regions across visual cortex. Our method performs evolutionary search over a structured prompt space, guided by a dynamic encoding model that predicts voxel-level responses to video inputs. By maximizing predicted activity for a target ROI, the framework efficiently discovers hyper-activating dynamic stimuli that consistently surpass handcrafted localizer videos. The synthesized videos recover known selectivities across ventral, dorsal, and lateral pathways, and further reveal systematic differences in sensitivity to temporal dynamics. A searchlight analysis provides new insight into the progression toward increasingly complex social-dynamic features along the lateral stream, further supported by probing with synthesized abstract, non-naturalistic stimuli. Taken together, our framework enables in silico exploration of dynamic visual selectivity, with new predictions for in vivo experiments
Comments10 pages, 6 figures