发表机构
University of Tehran(德黑兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现代视频变换器忽视灵长类视觉原理及缺乏神经数据评估的问题,提出含稀疏胜者全得令牌选择模块的脑对齐多流视频变换器,经实验其在精度与推理时间上表现优且鲁棒性强,脑模型相关性高,优于基线。
AI 中文摘要
现代视频变换器通常忽略灵长类视觉原理,且很少根据神经数据进行评估,限制了其生物学可解释性。我们引入了一个稀疏胜者全得令牌选择模块,取代密集自注意力以提高效率,并近似生物视觉回路中观察到的竞争路由。还提出了一种受神经启发的拆分融合视频变换器,它使用两个互补路径:高分辨率、低帧率的“什么”流和低分辨率、高帧率的“哪里”流,在分类前融合。在Kinetics-400和Something-Something V2上,我们最好的变体在可比规模和预训练的模型中处于精度与推理时间的帕累托前沿,对空间扰动的鲁棒性有所提高。通过对相同视频刺激的模型嵌入和时间分辨脑电图记录进行表征相似性分析,我们的模型达到了0.18的峰值脑模型相关性,持续优于强大的视频变换器基线,表明路径专业化和稀疏竞争是有效、脑对齐视频理解的有用归纳偏差。
英文摘要
Modern video transformers typically ignore principles from primate vision and are rarely evaluated against neural data, limiting their biological interpretability. We introduce a sparse winner-takes-all token selection module that replaces dense self-attention to improve efficiency and approximate competitive routing observed in biological visual circuits. We further propose a neuro-inspired split-and-fuse video transformer which uses two complementary pathways: a high-resolution, low-frame-rate "what" stream and a low-resolution, high-frame-rate "where" stream, fused before classification. On Kinetics-400 and Something-Something V2, our best variant operates on the Pareto frontier of accuracy versus inference time among models of comparable scale and pretraining, and showing improved robustness to spatial perturbations. Using representational similarity analysis between model embeddings and time-resolved EEG recordings for the same video stimuli, our model attains a peak brain-model correlation of 0.18 (about 78% of the noise ceiling) and consistently outperforms strong video transformer baselines, suggesting that pathway specialization and sparse competition are useful inductive biases for efficient, brain-aligned video understanding.
Comments23 pages, 11 figures