跨深度聚合与软路由专家在脑电基础模型微调中的数据集依赖效应
Dataset-Dependent Effects of Cross-Depth Aggregation and Soft-Routed Experts in EEG Foundation Model Fine-Tuning
- Vanderbilt University(范德堡大学)
- Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究在脑电基础模型微调中测试跨深度注意力残差与软路由专家模块,发现其效果因数据集而异,且带来显著计算开销,未产生一致增益。
中文摘要 AI 辅助
脑电(EEG)解码任务可能依赖于不同的时间动态和跨通道关系。我们测试了专用模块是否通过向CBraMod添加跨深度注意力残差(AttnRes)和两个软路由专家库来改进完全微调的脑电基础模型。在FACED、ISRUC、SEED-V和PhysioNet-MI上匹配的三种子实验中,完整模型相对于完全微调的平均平衡准确率变化分别为-0.12、+1.27、+0.77和-1.27个百分点。仅AttnRes在三个数据集上提高了平均平衡准确率,而在AttnRes之上添加专家仅对FACED和SEED-V有帮助。这些增益伴随着大量开销:AttnRes需要2.11至2.88倍的运行时间和1.78至2.67倍的内存,而完整模型需要2.41至3.04倍的运行时间和1.86至2.85倍的内存。总体而言,添加的模块产生了数据集依赖的、有时相互对立的效果,而非相对于完全微调的一致增益。
英文摘要
EEG decoding tasks can rely on different temporal dynamics and cross-channel relationships. We test whether specialized modules improve a fully fine-tuned EEG foundation model by augmenting CBraMod with cross-depth Attention Residuals (AttnRes) and two soft-routed expert banks. Across matched three-seed experiments on FACED, ISRUC, SEED-V, and PhysioNet-MI, the complete model changes mean balanced accuracy relative to full fine-tuning by -0.12, +1.27, +0.77, and -1.27 points, respectively. AttnRes alone improves mean balanced accuracy on three datasets, whereas adding experts on top of AttnRes helps only FACED and SEED-V. These gains come with substantial overhead: AttnRes requires 2.11 to 2.88x runtime and 1.78 to 2.67x memory, while the complete model requires 2.41 to 3.04x runtime and 1.86 to 2.85x memory. Overall, the added modules produce dataset-dependent, sometimes opposing effects rather than consistent gains over full fine-tuning.