arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27940cs.LGcs.CL

TriShield:通过正交梯度投影与优化器状态纠缠实现联邦语言模型微调中隐私后门的零效用损失防御

TriShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement

Cheng Wei

首次发表
浏览论文内容

中文总结 AI 辅助

TriShield是一种三层确定性防御,通过参数特征检测、有状态虚拟迭代和零效用正交投影,可完全抵御NeuroImprint攻击,且零模型效用损失、无额外通信轮次,计算开销低。

中文摘要 AI 辅助

大型语言模型(LLM)的联邦微调支持在不暴露原始数据的情况下开展协作训练。然而,近期一项名为NeuroImprint [1](arXiv:2606.20553)的攻击表明,恶意参数服务器可将PEFT适配器篡改隐私后门:通过为每个训练样本分配专用记忆神经元,并确保每个神经元最多更新一次,该服务器可高语义保真度地解析重构59%至79%的客户端训练数据。现有防御措施——包括本地差分隐私(LDP)[8]和梯度裁剪——要么无法抵御此类攻击,要么会导致不可接受的效用下降。本文提出TriShield,这是一种三层确定性防御机制,可在**零模型效用损失**且**无额外通信轮次**的前提下完全阻止NeuroImprint式重构。TriShield包含三个核心模块:(1)**参数特征检测器**,在本地训练开始前识别分布式模型参数中的记忆神经元特征;(2)**有状态虚拟迭代**机制,强制Adam/AdamW的动量状态在虚拟步骤间不可逆地纠缠梯度,使NeuroImprint的闭式反演失效;(3)**零效用正交投影**算子,将所有本地梯度更新投影到通过SVD计算的主任务语义子空间,物理消除任何携带私有记忆的梯度分量。我们从理论上证明,在第2和第3层处理后,上传梯度与任意单个训练样本间的互信息为零。在GPT-2(117M)和Llama-Guard-3-1B上开展的实验验证,TriShield可将所有测试攻击变体的NeuroImprint重构率降至**0%**,同时维持或提升训练准确率,且GPU计算额外开销不足5%。

英文摘要

Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data. However, a recent attack, NeuroImprint, demonstrates that a malicious parameter server can corrupt a PEFT adapter into a privacy backdoor: by assigning a dedicated memorization neuron to each training sample and ensuring each neuron updates at most once, the server can analytically reconstruct 59%--79% of client training data with high semantic fidelity. Existing defenses---including local differential privacy (LDP) and gradient clipping---either fail against this attack or impose unacceptable utility degradation. We present \textbf{TriShield}, a three-layer deterministic defense that completely prevents NeuroImprint-style reconstruction with zero model utility loss and no additional communication rounds. TriShield consists of: (1) a Parameter Artifact Detector that identifies memory-neuron signatures in distributed model parameters before local training begins; (2) a Stateful Virtual Iteration} mechanism that forces Adam/AdamW's momentum state to irreversibly entangle gradients across virtual steps, invalidating NeuroImprint's closed-form inversion; and (3) a Zero-Utility Orthogonal Projection operator that projects all local gradient updates onto the main-task semantic subspace computed via SVD, physically eliminating any gradient components that carry private memorization. We prove theoretically that after Layers 2 and 3, the mutual information between the uploaded gradient and any individual training sample is zero. Experiments on GPT-2 (117M) and Llama-Guard-3-1B verify that TriShield reduces NeuroImprint reconstruction rate to 0% across all tested attack variants, while maintaining or improving training accuracy, with less than 5% additional GPU computation overhead.

发表机构

  • Honor Device Co., Ltd.(荣耀终端有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑