使用端到端深度学习对视频刺激的脑电信号进行视觉语义解码
Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning
- The University of Melbourne(墨尔本大学)
- Bionics Institute(仿生研究所)
- The University of Osaka(大阪大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究利用端到端深度学习框架对视频刺激的脑电信号进行视觉语义解码,在每个视觉类别训练样本少的情况下评估多种方法,分析最佳方法依赖的信息,所选系统有良好性能且模型行为可解释,与神经科学知识相符。
AI中文摘要:
基于脑电信号(ECoG)的视觉语义解码能够从复杂、有噪声的大脑活动中推断视觉感知的语义解释。本研究探讨了使用端到端深度学习框架进行视觉语义解码的可行性。具体而言,解码任务是利用时间序列神经输入从视频刺激中预测视觉类别。使用先前收集的来自17名耐药性癫痫患者的ECoG数据集进行分析。每个视觉类别训练样本少于50个,评估了多种深度学习方法、人工神经网络架构和频带滤波输入。分析了性能最佳的方法所依赖的判别信息。所选解码系统使用混合增强、基于Transformer的编码器以及刺激后900毫秒的高伽马(80 - 150赫兹)输入。进一步分析表明,早期视觉皮层(V2 - V4)、腹侧流视觉皮层、MT +复合体及相邻视觉区域和外侧颞叶皮层对解码性能有显著贡献。本研究表明,端到端深度学习框架无需手工特征就能从动态视觉刺激中产生有前景的解码性能,且模型行为可通过频谱、时间和皮层维度进行解释,与现有神经科学知识大致一致。
英文摘要:
ECoG-based visual semantic decoding enables inference of semantic interpretation of visual perception from complex, noisy brain activity. This study examines the feasibility of visual semantic decoding using an end-to-end deep learning framework using electrocorticography (ECoG). Specifically, the decoding task is to predict visual categories from video stimuli using time-series neural inputs. A previously collected ECoG dataset from participants ($n=17$) with drug-resistant epilepsy is used for analysis. With fewer than 50 training samples per visual category, this study evaluates multiple deep learning approaches, artificial neural network architectures, and frequency-band filtered inputs. The best-performing approach is analyzed to shed light on the discriminative information it relies on across spectral, temporal, and cortical dimensions. The selected decoding system uses mixup augmentation, a Transformer-based encoder, and high-gamma (80-150 Hz) inputs with a 900 ms post-stimulus window. Further analysis shows that early visual cortex (V2-V4), ventral stream visual cortex, MT+ complex with neighbouring visual areas, and lateral temporal cortex contributed substantially to decoding performance. This study demonstrates that an end-to-end deep learning framework can yield promising decoding performance from dynamic visual stimuli without handcrafted features, while the model behavior remains interpretable through spectral, temporal, and cortical dimensions, which are broadly consistent with established neuroscience knowledge.