发表机构
The University of Electro-Communications; Artificial Intelligence eXploration Research Center (AIX)(电气通信大学; 人工智能探索研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对脑电到语音解码的低信噪比与会话变异性问题,采用多会话同一刺激的EEG响应作为对比学习正样本,结合变分正则化,在日语EEG数据集上降低了字符错误率并保持重构保真度。
AI 中文摘要
从无创脑电图(EEG)重构听到的语音极具挑战性,因为其信噪比(SNR)低且存在会话间变异性。虽然试次平均可提高信噪比,但难以应用于连续语音。我们将不同会话中同一刺激对应的重复EEG响应作为对比学习的正样本对,并引入变分正则化,结合该对比目标使编码器表示空间保持宽广。对日语EEG数据集的实验表明,会话不变策略与变分正则化结合可降低字符错误率(CER),同时保持梅尔频谱图重构保真度。会话探测证实编码器表示实现了会话不变性。
英文摘要
Reconstructing heard speech from non-invasive electroencephalography (EEG) is challenging due to a low signal-to-noise ratio (SNR) and inter-session variability. While trial averaging improves the SNR, it is difficult to apply to continuous speech. We instead use repeated EEG responses to the same stimulus across different sessions as positive pairs for contrastive learning, and introduce variational regularization that, combined with this contrastive objective, keeps the encoder representation space broad. Experiments on a Japanese EEG dataset show that combining the session-invariant strategy with variational regularization improves the character error rate (CER) while maintaining mel-spectrogram reconstruction fidelity. Session probing confirms that the encoder representations achieve session-invariance.
CommentsAccepted to APSIPA ASC 2026