发表机构
School of Intelligence Science and Technology, Peking University; Academy for Advanced Interdisciplinary Studies, Peking University; National Key Laboratory of General Artificial Intelligence(北京大学智能科学与技术学院; 北京大学前沿交叉学科研究院; 通用人工智能全国重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出SICMD方法,融合fMRI与MEG信号实现跨受试者感知语音解码,显著提升Top-1、Top-10和Rankacc指标,并大幅降低训练成本。
AI 中文摘要
基于非侵入式脑机接口(BCI)信号的感知语音解码近年来得到了广泛研究。该领域的研究主要面临两大挑战:提取具有丰富时空信息的神经表征以及实现跨受试者泛化。尽管已有独立研究提出了应对这些问题的不同方法,但同时解决这两大挑战的统一方法仍然缺乏。为填补这一空白,我们提出了主题不变跨模态感知语音解码(SICMD)方法,该方法整合了功能性磁共振成像(fMRI)和脑磁图(MEG)数据。我们对融合方法、融合位置、编码器架构以及模型输入进行了全面分析。结果表明,在跨受试者感知语音解码任务中,与基线方法相比,所提出的方法在Top-1、Top-10和Rankacc指标上分别提升了超过10.6%、10.1%和1.7%,同时与多受试者和单受试者解码设置相比,训练成本分别降低了88.8%和60.5%。进一步的视觉化实验也证实了我们方法的有效性。
英文摘要
Perceived speech decoding based on non-invasive brain-computer interface (BCI) signals has been extensively studied in recent years. Research in this field primarily faces two challenges: extracting neural representations with rich spatiotemporal information and achieving cross-subject generalization. Although separate studies have proposed methods to cope with these issues, a unified approach that simultaneously tackles both challenges remains lacking. To fill this gap, we propose the Subject-Invariant Cross-Modal Perceived Speech Decoding (SICMD) method, which integrates functional magnetic resonance imaging (fMRI) and magnetoencephalography (MEG). We conduct comprehensive analyses of the fusion method, fusion position, encoder architecture, and model inputs. Our results demonstrate that the proposed method improves Top-1, Top-10, and Rankacc by more than 10.6%, 10.1%, and 1.7%, respectively, compared to baseline methods in cross-subject perceived speech decoding tasks, while reducing training costs by 88.8% and 60.5% compared to multi-subject and intra-subject decoding settings. Further visualization experiments also confirm the effectiveness of our approach.
CommentsSubmitted to ICASSP 2027