发表机构
BITS Pilani(比拉理工学院皮拉尼校区)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在THINGS-EEG2上提出单被试EEG视觉刺激检索基线,使用时空卷积编码器映射到ViT特征,取得高于随机的语义解码,但直接重建像素仍不可行。
AI 中文摘要
从脑电图(EEG)重建视觉刺激是困难的,因为头皮测量具有高时间分辨率但空间分辨率有限,且配对的EEG-图像数据集相对于现代生成模型训练语料库仍然较小。我们在THINGS-EEG2上提出了一个可复现的单被试基线,首先测试了更有依据的问题,即EEG是否能在视觉嵌入空间中检索所观看的刺激。一个紧凑的时空卷积编码器将重复平均的EEG(63×250)映射到提供的512维ViT-B/32图像特征。模型选择使用概念不相交的验证集划分,最终评估使用官方的200图像、200概念测试画廊。在三个训练种子中,模型在1、5和10个检索项上分别获得12.83±0.58%、39.17±1.76%和58.00±1.73%的图像召回率(均值±样本标准差),而分析性随机水平分别为0.5%、2.5%和5.0%。一个会话平衡的消融实验表明,平均更多的测试重复通常改善排序。将Subject 01模型直接应用于其他九个被试而不进行适应会导致性能急剧下降,揭示了被试特异性。我们进一步报告了直接条件生成器的探索性压力测试,这些生成器在没有外部视觉权重的情况下训练:单被试和十被试变体产生噪声主导的输出,早期验证改进在一到四个epoch后逆转。最后,我们区分了直接重建与使用预训练扩散先验的语义渲染。结果支持在闭集、重复平均协议下高于随机水平的粗略语义解码,但不支持对刺激像素的忠实恢复。
英文摘要
Reconstructing visual stimuli from electroencephalography (EEG) is difficult because scalp measurements have high temporal but limited spatial resolution, and paired EEG-image datasets remain small relative to modern generative-model training corpora. We present a reproducible single-subject baseline on THINGS-EEG2 that first tests the more defensible question of whether EEG can retrieve the viewed stimulus in a visual embedding space. A compact temporal-spatial convolutional encoder maps repetition-averaged EEG (63 by 250) to provided 512-dimensional ViT-B/32 image features. Model selection uses a concept-disjoint validation split, and final evaluation uses the official 200-image, 200-concept test gallery. Across three training seeds, the model obtains 12.83 +/- 0.58%, 39.17 +/- 1.76%, and 58.00 +/- 1.73% image recall at 1, 5, and 10 (mean +/- sample standard deviation), compared with analytical chance levels of 0.5%, 2.5%, and 5.0%. A session-balanced ablation shows that averaging more test repetitions generally improves ranking. Applying the Subject 01 model to the other nine subjects without adaptation causes a sharp performance drop, exposing subject specificity. We further report exploratory stress tests of direct conditional generators trained without external visual weights: single-subject and ten-subject variants produce noise-dominated outputs, with early validation improvements reversing after one to four epochs. Finally, we distinguish direct reconstruction from semantic rendering with a pretrained diffusion prior. The results support above-chance coarse semantic decoding under a closed-set, repetition-averaged protocol, but do not support faithful recovery of stimulus pixels.