发表机构
Brown University; Stanford University; National University of Singapore; Zhejiang University; University of Oxford; Nanyang Technological University; Institute of Automation, Chinese Academy of Sciences(布朗大学; 斯坦福大学; 新加坡国立大学; 浙江大学; 牛津大学; 南洋理工大学; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ZeroMAG提出零样本多模态适配器生成框架,通过函数约束潜在空间和模态-主体-任务条件,为冻结脑电图基础模型生成适配器权重,在六个目标数据集上平均提升平衡准确率7.22个百分点,接近监督适配性能。
AI 中文摘要
脑电图基础模型(EFMs)从大规模脑电图数据中捕获可复用的知识,而许多脑电图记录还包含伴随的生理信号,这些信号提供了超出仅脑电图接口的互补信息。挑战在于在保留这种预训练知识的同时,通过从未标记的目标数据推断出的适配,将EFM扩展到异构多模态记录。我们提出了ZeroMAG,一种零样本多模态适配器生成框架,它使用未标记的目标记录扩展冻结的脑电图编码器和预测头,无需目标标签或目标端优化。目标数据集在ZeroMAG流程中从所有模型训练和选择中保留出来。ZeroMAG围绕配置不变的适配器组织伴随模态,从未标记记录和任务上下文构建模态-主体-任务条件,并在从源适配器学习的函数约束潜在空间中生成适配器权重。在六个保留的目标数据集和三个EFM骨干网络上,ZeroMAG相比仅脑电图推理将平衡准确率提高了7.22个百分点,相比直接权重回归提高了4.89个百分点,同时平均距离监督多模态适配仅差0.50个百分点。消融研究进一步表明,从表示学习或条件生成中移除功能监督会降低生成适配器的性能,证实了这两个组成部分的贡献。
英文摘要
EEG foundation models (EFMs) capture reusable knowledge from large-scale EEG data, while many EEG recordings also include companion physiological signals that provide complementary information beyond the EEG-only interface. The challenge is to preserve this pretrained knowledge while extending the EFM to heterogeneous multimodal recordings through an adaptation inferred from unlabeled target data. We introduce ZeroMAG, a zero-shot multimodal adapter generation framework that extends a frozen EEG encoder and prediction head using unlabeled target recordings, without target labels or target-side optimization. The target datasets are held out from all model training and selection in the ZeroMAG pipeline. ZeroMAG organizes companion modalities around a configuration-invariant adapter, constructs a modality-subject-task condition from unlabeled recordings and task context, and generates adapter weights in a function-constrained latent space learned from source adapters. Across six held-out target datasets and three EFM backbones, ZeroMAG improves balanced accuracy by 7.22 percentage points over EEG-only inference and 4.89 points over direct weight regression, while coming within 0.50 points of supervised multimodal adaptation on average. Ablations further show that removing functional supervision from either representation learning or conditional generation degrades generated-adapter performance, confirming the contribution of both components.
Comments41 pages