arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

交叉注意力编码模型揭示人类高级视觉皮层中的动态时空路由

Cross-attention encoding models reveal dynamic spatiotemporal routing across human higher visual cortex

Iishaan Inabathini, Margaret M. Henderson

arXiv 2609.36366首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于交叉注意力的视频编码模型,通过逐脑区动态路由V-JEPA-2特征,提升fMRI预测并揭示高级视觉皮层中刺激依赖的时空注意力机制。

AI 中文摘要

理解大脑如何从随时间变化的自然输入中解析动作和事件是神经科学的核心挑战。近期研究利用深度神经网络(DNN)模型构建了刺激可计算的fMRI编码模型,用于预测单个体素对复杂自然视频的反应。然而,大多数视频可计算编码模型通过模型令牌的简单线性映射来预测反应,忽视了视频表示与神经反应之间共享的时空结构。最近的交叉注意力编码模型针对静态图像解决了这一局限性,实现了跨空间对图像内容进行灵活的刺激依赖性加权。在此,我们将这一框架扩展到自然视频,使用逐脑区交叉注意力从自监督视频模型(V-JEPA-2)中动态路由特征,跨越空间和时间,并将该模型拟合到对短视频片段的fMRI反应。我们比较了联合时空注意力与分解及选择性受限的替代方案,发现联合路由在高级视觉区域中提高了对留出视频的脑反应预测,最一致的是在与动态运动感知相关的侧向和背侧视觉区域。此外,我们的方法提供了可解释的、刺激特定的注意力图,这些图动态跟踪移动物体,揭示每个神经反应所贡献的位置和时间时刻。我们进一步表明,来自不同类别选择性网络(面部、身体、场景选择性)的脑区注意力图根据预期的语义选择性差异性地加权内容。总之,这项工作提供了一个新的计算框架,用于理解在动态视觉感知过程中,视觉信息如何被皮层群体自适应加权。

英文摘要

Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of video-computable encoding models predict responses using simple linear mappings from model tokens, overlooking the spatiotemporal structure shared by video representations and neural responses. Recent cross-attention encoding models address this limitation for static images, enabling flexible stimulus-dependent weighting of image content across space. Here, we extend this framework to naturalistic video, using per-parcel cross-attention to dynamically route features from a self-supervised video model (V-JEPA-2) across both space and time, fitting this model to fMRI responses to short video clips. We compare joint spatiotemporal attention with factorized and selectively constrained alternatives, and find that joint routing improves predictions of brain responses to held-out videos across higher visual regions, most consistently in lateral and dorsal visual areas associated with dynamic motion perception. Moreover, our method provides interpretable, stimulus-specific attention maps that dynamically follow moving objects, revealing which locations and temporal moments contribute to each neural response. We further show that attention maps from parcels in different category-selective networks (face-, body-, scene-selective) differentially weight content in accordance with expected semantic selectivity. Together, this work provides a new computational framework for understanding how visual information is adaptively weighted by cortical populations during dynamic visual perception.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑