差分隐私脑电图特征匿名化:临床神经生理学中的隐私-效用案例研究
Differentially Private EEG Feature Anonymization: A Privacy-Utility Case Study in Clinical Neurophysiology
浏览论文内容
中文总结 AI 辅助
本研究针对临床EEG特征匿名化,提出基于高斯和拉普拉斯扰动的受试者级差分隐私框架,评估三种部署场景,揭示隐私与效用间的权衡及小样本不平衡数据的挑战。
中文摘要 AI 辅助
临床脑电图(EEG)数据对于医疗保健研究和开发基于人工智能(AI)的临床决策支持系统具有重要价值,但EEG记录和衍生特征可能包含敏感的患者特定信息。当数据在临床和研究环境中被重用、分析或共享时,这会产生隐私风险。传统的匿名化方法对于高维生物医学信号往往不足,因为移除直接标识符并不一定能防止重新识别、链接或推断风险。同时,强隐私保护可能会扭曲临床相关的信号特征并降低数据效用。本文研究了使用高斯和拉普拉斯扰动来保护临床EEG衍生特征表示的受试者级差分隐私。所提出的框架考虑了三种部署场景:客户端匿名化、集中式服务器端匿名化和去中心化本地训练。在EEG预处理和特征提取之后,对所得的患者级EEG特征表示应用高斯和拉普拉斯扰动。拉普拉斯实验评估了实现的噪声尺度,而正式全向量校准所需的尺度则单独推导。使用统计效用度量和下游基于机器学习的效用检查来评估两种扰动的影响。结果表明,差分隐私扰动可以集成到EEG处理流程中,但所选择的机制、隐私参数和灵敏度校准强烈影响数据效用。该研究强调了基于DP的EEG特征匿名化中实际的隐私-效用权衡,以及在小型和不平衡的临床EEG数据集中保持下游效用的挑战。
英文摘要
Clinical electroencephalography (EEG) data are valuable for healthcare research and for developing artificial intelligence (AI)-based clinical decision-support systems, but EEG recordings and derived features may contain sensitive patient-specific information. This creates privacy risks when data are reused, analyzed, or shared across clinical and research environments. Conventional anonymization methods are often insufficient for high-dimensional biomedical signals, since removing direct identifiers does not necessarily prevent re-identification, linkage, or inference risks. At the same time, strong privacy protection may distort clinically relevant signal characteristics and reduce data utility. This paper studies subject-level differential privacy for protecting clinical EEG-derived feature representations using Gaussian and Laplace perturbations. The proposed framework considers three deployment scenarios: client-side anonymization, centralized server-side anonymization, and decentralized local training. Following EEG preprocessing and feature extraction, Gaussian and Laplace perturbations are applied to the resulting patient-level EEG feature representations. The Laplace experiments evaluate the implemented noise scales, while the scales required for formal full-vector calibration are derived separately. The effects of both perturbations are assessed using statistical utility measures and a downstream machine-learning-based utility check. The results show that differentially private perturbation can be integrated into EEG processing workflows, but the selected mechanism, privacy parameters, and sensitivity calibration strongly influence data utility. The study highlights the practical privacy-utility trade-off in DP-based EEG feature anonymization and the challenges of preserving downstream utility in small and imbalanced clinical EEG datasets.
发表机构
- University of South-Eastern Norway(东南挪威大学)
机构由 AI 辅助整理,请以论文原文为准。