2026 PNPL竞赛:LibriBrain100中的单词分类与高效跨主体泛化
The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100
浏览论文内容
中文总结 AI 辅助
2026 PNPL竞赛基于扩展的LibriBrain100数据集,设置单词分类的Deep与Broad赛道,以推进非侵入式BCI的跨主体泛化能力。
中文摘要 AI 辅助
2025 PNPL竞赛(Landau等人,2025)的目标是为非侵入式语音解码推出一个多年课程,旨在从基础任务逐步推进到实用脑机接口(BCI)所需的语言复杂度,以语音检测和音素分类任务拉开序幕。获胜提交在对应任务上达到了宏F1分数95.6%和73.6%(Elvers等人,2026),是极具意义的进展。这一成功基于LibriBrain数据集(Özdogan等人,2025),该数据集是当时记录的最大的主体内MEG数据集,单个主体拥有约50小时数据。然而,尽管主体内规模可带来强大的解码性能,实用BCI必须能从数分钟而非数小时的数据中泛化到新用户。2026 PNPL竞赛针对这一挑战推出了LibriBrain100(Mantegna等人,2026),这是扩展后的LibriBrain数据集,包含32个额外主体(每个约40分钟数据),以及更多的主体内数据(约80小时)。竞赛将任务课程推进至聚焦单词分类,设置了两个互补赛道:Deep赛道针对大规模主体内单词分类,旨在达到最佳性能;Broad赛道针对跨主体泛化,逐步将主体特定微调数据量从约40分钟降至约20分钟,再到约10分钟,该时长处于临床可行范围内,让我们更接近能为重度瘫痪患者恢复交流能力的非侵入式BCI。
英文摘要
The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-invasive speech decoding. Designed to progress from foundational tasks toward the linguistic complexity required for a practical brain-computer interface (BCI), it set the stage with speech detection and phoneme classification tasks. Winning submissions reached F1-macro scores of 95.6% and 73.6% on the respective tasks (Elvers et al., 2026), highly significant advances. This success was built on the LibriBrain dataset (Özdogan et al., 2025), the largest within-subject MEG dataset recorded at the time with ${\sim}50$ hours of data for one subject. However, while within-subject scale drives strong decoding performance, a practical BCI must generalise to new users from minutes of data, not hours. The 2026 PNPL competition responds to this challenge with LibriBrain100 (Mantegna et al., 2026), an extended LibriBrain dataset with 32 additional subjects (${\sim}40$ minutes each) plus even more within-subject data (${\sim}80$ hours). Advancing the curriculum of tasks to focus on word classification, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation, progressively reducing the amount of subject-specific fine-tuning data from ${\sim}40$ to ${\sim}20$ to ${\sim}10$ minutes, a duration that falls within a clinically feasible range and brings us a step closer to a non-invasive BCI capable of restoring communication to people living with profound paralysis.
发表机构
- University of Oxford(牛津大学)
- Maastricht University(马斯特里赫特大学)
机构由 AI 辅助整理,请以论文原文为准。