结构化视觉目标学习用于跨被试脑电-图像检索
Structured Visual Target Learning For Cross-Subject eeg-to-image retrieval
- Indian Institute of Technology Roorkee(印度理工学院鲁尔基分校)
- University of La Rochelle(拉罗谢尔大学)
- Indian Institute of Technology (ISM) Dhanbad(印度理工学院(ISM)丹巴德分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对跨被试脑电-图像检索,提出从视觉目标侧学习结构化表征,结合块路由与MMD正则,并引入免训练精炼,在THINGS-EEG2上显著提升检索精度。
AI中文摘要:
跨被试脑电-图像检索要求一个在源被试上训练的神经表征,对于未见过的被试,仍能与视觉嵌入空间保持对齐。现有方法主要关注脑电侧,而我们则从视觉目标的角度来解决这一问题。我们的方法保留了感知编码器的空间信息,将其补丁网格转换为一组紧凑的学习到的视觉视图,并通过一个块结构化的、内容相关的路由器为每张图像聚合这些视图。该目标与脑电编码器通过对比学习联合训练,并在源被试间使用最大均值差异(MMD)正则化。在部署时,我们提出了一种无需训练的表示精炼方法,在不更新任一编码器的情况下对齐冻结的嵌入。在THINGS-EEG2上的留一被试评估中,结构化目标实现了35.3%/65.6%的Top-1/Top-5准确率,在比较方法中表现最佳。精炼方法将其提升至48.1%/77.1%,相比最强比较方法,Top-1准确率提高了18.5%,并改善了所有十个留出被试的结果。
英文摘要:
Cross-subject EEG-to-image retrieval requires a neural represen- tation trained on source subjects to remain aligned with a visual embedding space for an unseen subject. Whereas existing methods primarily focus on the EEG side, we address this problem from the perspective of the visual target. Our approach preserves the spatial information of the Perception Encoder, converts its patch grid into a compact set of learned visual views, and aggregates them for each image with a block-structured, content-dependent router. The target is learned jointly with the EEG encoder through contrastive learning with MMD regularization across source subjects. For deployment, we propose a training-free representation refinement that aligns frozen embeddings without updating either encoder. Under leave- one-subject-out evaluation on THINGS-EEG2, the structured target achieves 35.3%/65.6% Top-1/Top-5 accuracy, the best among com- pared methods. Refinement raises this to 48.1%/77.1%, an 18.5% Top-1 gain over the strongest compared method, improving all ten held-out subjects.