SW-ProxyCE:从公开EEG编码器到私有下游模型的零查询对抗迁移
SW-ProxyCE: Zero-Query Adversarial Transfer from Public EEG Encoders to Private Downstream Models
浏览论文内容
中文总结 AI 辅助
本文针对公开编码器与私有下游模型场景,提出无查询的任务感知攻击框架SW-ProxyCE,实验表明其生成的对抗样本可有效迁移至下游模型,且性能优于任务无关攻击。
中文摘要 AI 辅助
脑电图(EEG)基础模型近来成为通过从大规模异构神经记录中学习可复用表征来进行EEG解码的有前景范式。然而,EEG基础编码器的公开发布在推动下游发展的同时,也引入了此前未被探索的安全风险:公开可用的表征可能使私有下游模型易受攻击。本文研究了公开编码器与私有下游模型场景下EEG基础模型部署中的对抗迁移攻击,其中攻击者可白盒访问已发布的编码器和小型任务匹配的带标签参考集,但无法访问或查询受害者的参数、输出或梯度。我们提出Shrinkage-Whitened Proxy Cross-Entropy(SW-ProxyCE),这是一种无查询的任务感知攻击框架,通过收缩白化类原型从小型带标签参考集中恢复任务级决策几何,无需训练额外的代理分类器即可实现可迁移的对抗样本生成。我们在三个EEG任务上使用三个通用基础编码器和一个特定范式的预训练编码器评估了SW-ProxyCE,涵盖跨被试和被试内场景下的线性探测和全微调下游模型。结果表明,从公开编码器和有限带标签参考生成的对抗样本可有效迁移到无法访问的下游模型,SW-ProxyCE始终优于任务无关的表征偏移攻击,揭示EEG基础模型的强可迁移性并不一定带来对抗鲁棒性。我们的代码将在GitHub上发布。
英文摘要
Electroencephalography (EEG) foundation models have recently emerged as a promising paradigm for EEG decoding by learning reusable representations from large-scale heterogeneous neural recordings. However, the open release of EEG foundation encoders, while facilitating downstream developments, also introduces a previously unexplored security risk: publicly available representations may make private downstream models vulnerable. This paper investigates adversarial transfer attacks in EEG foundation model deployment in a public-encoder and private-downstream setting, where attackers have white-box access to a released encoder and a small task-matched labeled reference set, but no access or query to victim parameters, outputs, or gradients. We propose Shrinkage-Whitened Proxy Cross-Entropy (SW-ProxyCE), a query-free task-aware attack framework that recovers task-level decision geometry from a small labeled reference set through shrinkage-whitened class prototypes, enabling transferable adversarial generation without training an additional surrogate classifier. We evaluated SW-ProxyCE across three EEG tasks using three general-purpose foundation encoders and a paradigm-specific pre-trained encoder, covering both linear-probing and full-fine-tuning downstream models in cross-subject and within-subject scenarios. Results demonstrated that adversarial examples generated from the public encoder and limited labeled references can effectively transfer to inaccessible downstream models. SW-ProxyCE consistently outperformed task-agnostic representation-shift attacks, revealing that the strong transferability of EEG foundation models does not necessarily lead to adversarial robustness. Our code will be available on GitHub.