arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19534cs.CR

AEGIS:注意力嵌入梯度隔离护盾——用于隐私保护的联邦大语言模型微调的三通道梯度掩码

AEGIS: Attention-Embedding Gradient Isolation Shield - Triple-Channel Gradient Masking for Privacy-Preserving Federated LLM Fine-Tuning

Ye Tao, Hong Shen, Hui Tian, Xin Wang, Can Wang

首次发表
浏览论文内容

中文总结 AI 辅助

AEGIS是一种轻量级联邦大语言模型微调防御措施,通过三通道梯度掩码关闭私有令牌信息泄露通道,可将梯度反演攻击的令牌恢复率降至接近零,同时保留或提升模型效用。

中文摘要 AI 辅助

梯度反演攻击可从联邦学习中共享的梯度恢复私有训练文本,对协作式模型训练构成严重威胁。通过对Transformer梯度结构的分析,我们确定了私有令牌信息泄露的三个通道:注意力输出投影梯度暴露了编码输入嵌入的低秩子空间(通道1);嵌入梯度的行范数稀疏性直接揭示了存在哪些令牌(通道2);MLP扩展梯度携带与通道1类似的可恢复子空间信号(通道3)。现有最先进的攻击方法会利用这些通道进行分析,在数秒内实现近乎精确的令牌恢复。现有的防御措施最多仅能应对一个通道,要么会降低模型效用,要么会保留其余结构信号。我们引入了AEGIS(注意力嵌入梯度隔离护盾),这是一种轻量级防御措施,通过三个反向路径操作关闭所有三个分析通道,且无需对架构进行任何更改:冻结注意力投影参数可从结构上消除通道1;对嵌入梯度注入校准噪声可破坏通道2的令牌存在信号;对MLP扩展梯度注入类似的每块噪声可掩码通道3。相同的掩码梯度同时用于本地优化器步骤和服务器导出,因此两侧均不会保留干净信号。在11个模型和6个数据集上进行评估后,AEGIS可将针对一系列梯度反演攻击(包括分析型和优化型)的令牌恢复率降至接近零,同时保留或提升模型效用。我们为通道1和通道2提供了形式化保证,并针对完全了解该机制的自适应对手对完整防御措施进行了实证验证。

英文摘要

Gradient inversion attacks recover private training text from gradients shared in federated learning, posing a serious threat to collaborative model training. Through our analysis of transformer gradient structure, we identify three channels through which private token information leaks: the attention output projection gradient exposes a low-rank subspace that encodes input embeddings (Channel 1), the embedding gradient's row-norm sparsity directly reveals which tokens are present (Channel 2), and the MLP expansion gradient carries a recoverable subspace signal analogous to Channel 1 (Channel 3). State-of-the-art attacks exploit these channels analytically to achieve near-exact token recovery in seconds. Existing defences address at most one channel and either degrade model utility or leave the remaining structural signals intact. We introduce AEGIS (Attention-Embedding Gradient Isolation Shield), a lightweight defence that closes all three analytical channels with three backward-path operations requiring no architectural changes: freezing attention projection parameters eliminates Channel 1 by construction, calibrated noise injection into the embedding gradient destroys Channel 2's token-presence signal, and analogous per-block noise injection into the MLP expansion gradient masks Channel 3. The same masked gradient drives both the local optimiser step and the server export, so no clean signal is retained on either side. Evaluated across 11 models and six datasets, AEGIS reduces token recovery rates to near zero against a range of gradient inversion attacks, both analytical and optimisation-based, while preserving or improving model utility. We provide formal guarantees for Channels 1 and 2 and validate the full defence empirically against adaptive adversaries with complete knowledge of the mechanism.

补充信息

↑