arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Cephalonauts One:用于解码人脑自然语音的深度fMRI数据集

Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain

Antoine Collas, Louis Jalouzot, Géraud Ilinca, Corentin Caris, Romain Valabrègue, Ahmed Hassayoune, David Goncalves, Madeleine Hueber, Thaddée Delebarre, Julien Savatovsky, Clara Fonteneau, Charles Maussion, Bertrand Thirion, Alexis Thual

arXiv 2610.03558首次发表:更新:

发表机构

Karavela; UNICOG, CNRS, INSERM, CEA, Paris-Saclay University; LSCP, EHESS, ENS, CNRS, PSL University; Centre de NeuroImagerie de Recherche (CENIR), Sorbonne Université, ICM; Department of Radiology, Hôpital Fondation Adolphe de Rothschild; Inria, CEA, Paris-Saclay University(Karavela; UNICOG,CNRS,INSERM,CEA,巴黎萨克雷大学; LSCP,EHESS,ENS,CNRS,巴黎文理研究大学; 索邦大学巴黎大脑研究所神经影像研究中心(CENIR); 罗斯柴尔德基金会医院放射科; 法国国家信息与自动化研究所,CEA,巴黎萨克雷大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发布了一个包含三名受试者各30小时fMRI数据的深度数据集,用于解码自然语音,并提出了音频片段检索基准,实验表明解码性能随训练数据量增加而提升。

AI 中文摘要

Cephalonauts One是一个全脑3特斯拉(3T)功能磁共振成像(fMRI)数据集,记录于受试者收听音频播客时。三名健康受试者接受了多次扫描会话,每次会话包含五个15分钟的运行,期间收听其母语的播客。每个受试者拥有30小时的fMRI数据,当前版本是使用自然语音刺激的最深可用的fMRI数据集。该数据集将大脑活动与相应的播客音频、转录注释和派生的刺激嵌入配对。此外,我们引入了一个以音频片段检索形式制定的脑解码基准:给定来自保留会话的fMRI活动,解码器必须在候选片段中识别出对应的时间对齐的播客音频片段。我们为此任务提供了标准化的分割、评估指标和基线解码器。最后,缩放分析显示,解码性能随着每个受试者的训练数据量增加而持续提高。

英文摘要

Cephalonauts One is a whole-brain 3 Tesla (3T) functional magnetic resonance imaging (fMRI) dataset recorded while subjects listened to audio podcasts. Three healthy subjects underwent multiple scanning sessions, each consisting of five 15-minute runs, while listening to podcasts in their native language. With 30 hours of fMRI data per subject, the current release is the deepest available fMRI dataset using naturalistic speech stimuli. The dataset pairs brain activity with the corresponding podcast audio, transcript annotations, and derived stimulus embeddings. Furthermore, we introduce a brain decoding benchmark formulated as audio segment retrieval: given fMRI activity from a held-out session, the decoder must identify the corresponding time-aligned podcast audio segment among candidate segments. We provide standardized splits, evaluation metrics, and baseline decoders for this task. Finally, a scaling analysis shows that decoding performance improves continuously with the amount of training data per subject.

CommentsAccepted at NeurIPS 2026, Evaluations & Datasets Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑