arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22420cs.SDcs.AI

从非侵入式脑记录中解码感知语音的跨被试泛化

Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings

Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出CPSD框架,结合对比学习、PESA模块等,在三类感知语音数据集上跨被试解码性能优于基线,提升了Top-10准确率。

中文摘要 AI 辅助

近年来,从非侵入式脑记录中解码感知语音因其广泛的潜在应用而受到大量关注。然而,现有方法在跨被试解码中面临诸多挑战,主要源于泛化能力有限,且缺乏提取被试一致信息的显式机制。这些限制导致训练成本高、解码性能欠佳。为应对这些挑战,我们提出创新的跨被试感知语音解码(Cross-Subject Perceived Speech Decoding, CPSD)框架,包含两个训练阶段:源模型预训练与个性化适配。在源模型预训练阶段,采用对比学习捕获多个源被试间的共享表征;随后,个性化适配阶段通过从源模型中提取一致组件并利用目标被试数据微调,为目标被试初始化模型。此外,我们引入基于位置编码的空间注意力(Positional Encoding-based Spatial Attention, PESA)模块,该模块将脑磁图(MEG)/脑电图(EEG)数据重映射至标准化参考空间,从而增强跨被试一致性并简化模型训练。我们在涵盖不同模态与语言的三个感知语音神经数据集上评估所提CPSD框架,结果显示,该框架在Armeni 2022、PKUEEG 2025和Broderick 2018数据集的Top-10准确率上分别优于基线方法6.8%以上、15.4%和15.8%。进一步分析证实了所提方法的有效性、效率与鲁棒性。

英文摘要

Decoding perceived speech from non-invasive brain recordings has garnered significant attention in recent years due to its wide range of potential applications. However, existing methods face considerable challenges in cross-subject decoding, primarily due to limited generalizability and the absence of explicit mechanisms for extracting subject-consistent information. These limitations result in high training costs and suboptimal decoding performance. To address these challenges, we propose an innovative Cross-Subject Perceived Speech Decoding (CPSD) framework, which comprises two training stages: source model pre-training and personal specialization. In the source model pre-training stage, contrastive learning is employed to capture shared representations across multiple source subjects. Subsequently, personal specialization initializes the model for the target subject by extracting consistent components from the source model and fine-tuning it using target subject data. Additionally, we introduce the Positional Encoding-based Spatial Attention (PESA) module, which remaps MEG/EEG data into a standardized reference space, thereby enhancing cross-subject consistency and facilitating model training. We evaluate the proposed CPSD framework on three perceived speech neural datasets encompassing different modalities and languages. The results demonstrate that our framework outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top-10 accuracy on the Armeni 2022, PKUEEG 2025, and Broderick 2018 datasets, respectively. Further analyses confirm the effectiveness, efficiency, and robustness of the proposed approach.

发表机构

  • School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
  • Academy for Advanced Interdisciplinary Studies, Peking University(北京大学前沿交叉学科研究院)
  • National Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑