超越并行跟踪:交互式多特征融合驱动非侵入性脑记录的语义重建
Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings
浏览论文内容
中文总结 AI 辅助
研究针对非侵入性脑记录语义重建中表征不匹配问题,引入多特征融合框架,通过交互式门控机制结合静态与动态表征,经实验对比线性连接和非线性交叉注意力等方法,证明交叉注意力融合性能最佳,提供了新的脑到文本解码方法。
中文摘要 AI 辅助
从非侵入性神经记录进行连续语义重建仍然受到语义特征空间与神经编码模式之间表征不匹配的限制,严重阻碍了高噪声神经信号与目标语义特征之间的跨模态对齐。先前的语义解码器主要单独依赖静态词汇表征或动态上下文表征。这种单维方法不可避免地导致严重信息丢失。为弥合差距,本研究引入用于非侵入性语义重建的多特征融合框架,系统地对线性朴素连接和非线性多头交叉注意力这两种整合方法进行基准测试。通过广泛的语义重建和文本生成实验,揭示了强大的性能层次:交叉注意力>连接>GPT>词向量。关键的是,非线性交叉注意力融合方法实现了最先进的性能,证明神经语言解码受益于模拟上下文信息与核心词汇属性之间的协作调制,而非依赖孤立的个体特征,还提供了可行的非侵入性脑到文本解码方法。
英文摘要
Continuous semantic reconstruction from non-invasive neural recordings remains limited by the representational mismatch between semantic feature spaces and neural coding patterns, which severely impedes cross-modal alignment between high-noise neural signals and target semantic features. Prior semantic decoders have predominantly relied on static lexical representations or dynamic contextualized representations in isolation. This single-dimension approach inevitably leads to severe information loss, as it fails to account for the human brain's capacity to integrate stable word attributes and dynamic contexts simultaneously. To bridge this gap, this study introduces a multi-feature fusion framework for non-invasive semantic reconstruction, systematically benchmarking two integration approaches: linear Naive Concatenation and non-linear Multi-Head Cross-Attention. Within this framework, our approach complements static lexical representations (W2V) with dynamic contextual representations (GPT) via an interactive gating mechanism to facilitate cooperative processing during language comprehension. Evaluated through extensive semantic reconstruction and text generation experiments, our framework reveals a robust performance hierarchy: Cross-Att > Concat > GPT > W2V. Crucially, the non-linear cross-attention fusion method achieves state-of-the-art performance, demonstrating that neural language decoding benefits from simulating the collaborative modulation between contextual information and core lexical attributes rather than depending on isolated individual features, while also offering a viable non-invasive brain-to-text decoding method.
发表机构
- Center for BioMed-X Research, Academy for Advanced Interdisciplinary Studies,Peking University(北京大学生物医学前沿创新中心、前沿交叉学科研究院)
- Speech and Hearing Research Center, School of Intelligence Science and Technology,Peking University(北京大学智能科学与技术学院言语听觉研究中心)
- State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)
- National Biomedical Imaging Center, State Key Laboratory of Membrane Biology, Institute of Molecular Medicine, Peking-Tsinghua Center for Life Sciences, College of Future Technology, Peking University(国家生物医学成像中心、膜生物学国家重点实验室、分子医学研究所、北京大学生命科学联合中心、未来技术学院)
机构由 AI 辅助整理,请以论文原文为准。