发表机构
Chongqing University of Posts and Telecommunications; Shanghai Jiaotong University; Beijing University of Chemical Technology; The Hong Kong Polytechnic University(重庆邮电大学; 上海交通大学; 北京化工大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对脑到图像检索中固定最终层视觉目标的问题,提出NeuroGlyph,从多深度视觉骨干学习因子化目标并分配深度权重,在THINGS-EEG/MEG上超越基线。
AI 中文摘要
脑到图像检索旨在识别引发非侵入性神经反应的视觉刺激。候选图像通常由预训练视觉模型表示,其内部表示在深度上抽象程度不同。现有方法通常训练神经编码器以恢复固定的最终层视觉目标。在这种公式下,视觉层级被简化为一个规定的端点,阻止其他深度的表示直接塑造视觉目标。这一局限性促使我们学习视觉深度信息应如何贡献于检索目标。为此,我们引入了NeuroGlyph,它从冻结视觉骨干的多个深度学习一个试验无关的视觉目标。NeuroGlyph将目标分解为因子特定子空间。每个子空间学习一个基于图像条件的视觉深度分配。所得子空间被融合为用于检索的单一嵌入。在THINGS-EEG和THINGS-MEG上,NeuroGlyph在所有受控比较中均优于最终层监督。它还在四项比较中的三项超越了事后最佳固定层预言机。参数匹配的消融实验支持因子化目标构建和基于图像条件的深度分配。在可比较的200路检索协议下,NeuroGlyph在八项报告指标中的六项实现了最强的系统级性能。这些结果支持跨视觉层级学习检索目标,而非规定单一视觉深度。
英文摘要
Brain-to-image retrieval seeks to identify the visual stimulus that elicited a non-invasive neural response. Candidate images are typically represented by pretrained vision models, whose internal representations vary in abstraction across depth. Existing methods usually train the neural encoder to recover a fixed final-layer visual target. Under this formulation, the visual hierarchy is reduced to a single prescribed endpoint, preventing representations at other depths from directly shaping the visual target. This limitation motivates learning how information across visual depths should contribute to the retrieval target. To this end, we introduce NeuroGlyph, which learns a trial-independent visual target from multiple depths of a frozen visual backbone. NeuroGlyph decomposes the target into factor-specific subspaces. Each subspace learns an image-conditioned allocation over visual depth. The resulting subspaces are fused into a single embedding for retrieval. Across THINGS-EEG and THINGS-MEG, NeuroGlyph outperforms final-layer supervision in all controlled comparisons. It also surpasses the post hoc best fixed-layer oracle in three of four comparisons. Parameter-matched ablations support both factorized target construction and image-conditioned depth allocation. Under comparable 200-way retrieval protocols, NeuroGlyph achieves the strongest system-level performance in six of eight reported metrics. These results support learning retrieval targets across the visual hierarchy rather than prescribing one visual depth.