基于脑电图的视觉检索与重建:从神经可见最优层到分层扩散生成
EEG-based Visual Retrieval and Reconstruction: From Neurally Visible Optimal Layer to Hierarchical Diffusion Generation
浏览论文内容
中文总结 AI 辅助
本研究针对现有EEG视觉解码的跨模态不匹配问题,提出基于神经可见最优层(NVOL)的分层框架,结合对比学习与条件扩散,在THINGS-EEG数据集上提升了视觉检索准确率与图像重建效果。
中文摘要 AI 辅助
从脑电图(EEG)解码视觉感知对非侵入式脑机接口(BCI)具有重要意义。然而,大多数现有的视觉解码流程直接将EEG特征与预训练视觉模型的语义特征对齐,这些EEG信号携带多个层面的信息,而该做法忽略了EEG信号中不同视觉成分的 varying 神经可见性,导致跨模态不匹配和信息利用不充分。在本研究中,我们通过分层对比学习解决这一局限。针对每个受试者,选择使检索性能最大化的中间CLIP层作为神经可见最优层(NVOL)。基于NVOL,一个分层框架通过共享中间表示耦合检索与生成。检索分支融合多NVOL特征,通过对比学习将其与图像嵌入对齐,并在测试时应用跨域相似度局部缩放(CSLS)以缓解中心性问题。生成分支使用条件扩散先验从EEG重建受试者特定的NVOL特征,通过轻量适配器将其映射到CLIP空间,并驱动预训练的Stable Diffusion XL模型。在THINGS-EEG上的实验验证表明,基于NVOL的检索在200路检索中达到78.1%的平均Top-1准确率,应用CSLS后升至86.4%。两阶段NVOL到语义的重建在语义和结构指标上也优于单阶段最终层扩散。通过将EEG与分层神经可见性而非固定高级语义对齐,所提出的框架提升了基于EEG的视觉解码中的检索准确率和图像重建效果。
英文摘要
Decoding visual perception from electroencephalography (EEG) is important for non-invasive brain-computer interfaces (BCIs). However, most existing visual decoding pipelines directly align EEG features with semantic features from pretrained vision models. Those EEG signals carry information at more than one level and this practice disregards the varying neural visibility of different visual components in EEG signals, leading to cross modal mismatches and incomplete information use. In this work, we address this limitation through layer-wise contrastive learning. For each subject, the intermediate CLIP layer that maximizes retrieval performance is selected as the Neural Visibility Optimal Layer (NVOL). Built on NVOL, a hierarchical framework couples retrieval and generation through a shared intermediate representation. The retrieval branch fuses multi-NVOL features, aligns them to image embeddings via contrastive learning, and applies cross-domain similarity local scaling (CSLS) at test time to mitigate hubness. The generation branch reconstructs subject-specific NVOL features from EEG using a conditional diffusion prior, maps them to CLIP space through a lightweight adapter, and drives a pretrained Stable Diffusion XL model. Experimental validation on THINGS-EEG showed that, NVOL-based retrieval achieves 78.1\% mean Top-1 accuracy in 200-way retrieval, rising to 86.4\% with CSLS. Two-stage NVOL-to-semantic reconstruction also outperforms single-stage final-layer diffusion on semantic and structural metrics. By aligning EEG with layer-wise neural visibility rather than fixed high-level semantics, the proposed framework improves both retrieval accuracy and image reconstruction in EEG-based visual decoding.
发表机构
- University of Macau(澳门大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。