发表机构
HSE University; ITMO University(高等经济大学; ITMO大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究重新设计MEG到音频检索架构,在MEG-MASC数据集上实现39.75±0.34%的Top-1准确率,通过源映射与输入干预明确了驱动语音检索的关键刺激特征。
AI 中文摘要
通过使用CLIP风格目标函数针对wav2vec 2.0音频嵌入训练的深度网络,可从非侵入式脑磁图(MEG)记录中检索到感知语音的短片段,但这些网络的权重并未映射到电生理量,且仍不清楚哪些语音属性驱动了检索。我们在高性能MEG到音频检索架构的基础上,重新设计了其前端和解码器:原空间注意力作用于扁平化传感器布局,我们将其替换为基于三维MEG头盔几何结构定义的球谐函数;将被试特异性表征从270个分支缩减至25个,为每个分支添加时间滤波器以匹配时空上的神经元源,并将卷积解码器设计得更浅;在训练前去除眼动和心脏成分,以降低刺激锁定捷径的风险。在MEG-MASC数据集上,该模型在6种训练方案的1005个候选样本中达到39.75±0.34%的Top-1准确率,解码器参数数量减少约20倍。其权重映射到源空间,恢复出与语音感知网络一致的发生器,而左侧偏侧化分支携带右侧未显现的高频节律成分。配对MEG遮挡实验显示,19个刺激特征中有15个起作用,其中最大效应来自静音、声强、元音和声学起始;随机词列表则表现相反:将叙事性MEG替换入随机词列表会提升检索效果,表明无叙事结构的活动比连贯语音期间的活动携带更少可检索信息。wav2vec目标可缩减至约12个学习特征维度而准确率无损失,而强时间压缩则会导致明显损失。综上,源映射与输入干预揭示了驱动检索的因素。
英文摘要
Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval. We build on a high-performing MEG-to-audio retrieval architecture but redesign both its front end and decoder. Its spatial attention operates on a flattened sensor layout; we replace it with spherical harmonics defined on the three-dimensional MEG helmet geometry. We reduce the subject-specific representation from 270 to 25 branches, add a temporal filter to each branch to match it to a neuronal source in space and time, and make the convolutional decoder shallower. Ocular and cardiac components are removed before training to reduce the risk of stimulus-locked shortcuts. On MEG-MASC, the model reaches 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions, with about 20 times fewer decoder parameters. Its weights map to source space, recovering generators consistent with the speech-perception network, while left-lateralized branches carry higher-frequency rhythmic components not evident on the right. Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets. Random word lists behave oppositely: substituting narrative MEG into them improves retrieval, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target can be reduced to about twelve learned feature dimensions without loss of accuracy, whereas strong temporal compression causes a clear loss. Together, source mapping and input interventions reveal what drives retrieval.