arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27836cs.CRcs.AIcs.DC

FISGuard:通过固定输入子空间防御成员推理攻击

FISGuard: Defending Against Membership Inference via Fixed Input Subspaces

Haocheng Jiang, Hua Shen

首次发表
浏览论文内容

中文总结 AI 辅助

针对ProjRes利用梯度几何结构的成员推理攻击,本文提出FISGuard防御方法,通过固定公开数据构建的低维表示子空间,在保持模型效用的同时有效降低攻击效果,实现了隐私与效用的良好平衡。

中文摘要 AI 辅助

随着大语言模型在联邦学习中日益普及,在分布式私有数据上进行参数高效微调的同时保护用户隐私已成为重要挑战。尽管客户端仅共享梯度而非直接上传原始数据,但共享梯度仍可能泄露训练样本的成员信息。ProjRes(发表于S&P 2026)进一步加剧了这一风险:攻击者仅需候选表示与服务器可观测梯度诱导的子空间之间的投影残差,无需访问模型输出或获取更多信息,即可有效区分成员样本与非成员样本。现有针对成员推理攻击的防御方法大多依赖梯度扰动或正则化,这些方法不仅会降低模型效用,还无法有效抵御利用梯度几何结构的ProjRes成员推理攻击。为解决该问题,本文提出一种轻量级防御方法FISGuard,其核心思路是利用独立的公开数据构建并固定低维表示子空间,从而限制私有表示通过梯度暴露的空间,同时保留下游任务所需的核心信息,大幅缩小成员与非成员样本间的投影残差差异。我们在三个NLP数据集、两个大语言模型(LLM)及Adapter、LoRA两种微调策略上,将FISGuard与五种代表性防御方法进行评估。结果显示,在多数设置下,FISGuard可将ProjRes攻击的AUC降至接近随机猜测的0.5水平,同时保持下游任务性能与未防御模型相近,仅引入有限的计算开销,实现了良好的隐私-效用权衡。

英文摘要

As large language models are increasingly adopted in federated learning, protecting user privacy while performing parameter-efficient fine-tuning on distributed private data has become an important challenge. Although clients only share gradients instead of directly uploading raw data, the shared gradients may still leak membership information about training samples. ProjRes (S&P, 2026) further increases this risk: with less information and without accessing model outputs, an attacker can effectively distinguish members from non-members solely based on the projection residual between a candidate representation and the subspace induced by server-observable gradients. Existing defenses against membership inference mostly rely on gradient perturbation or regularization, which can not only degrade model utility but also fail to effectively defend against the membership inference attack introduced by ProjRes, which exploits the geometric structure of gradients. To address this issue, we propose FISGuard, a lightweight defense. Its key idea is to construct and fix a low-dimensional representation subspace using independent public data, thereby restricting the space through which private representations are exposed via gradients while preserving the primary information required for downstream tasks. This substantially reduces the projection-residual discrepancy between members and non-members. We evaluate FISGuard against five representative defense methods across three NLP datasets, two LLMs, and two fine-tuning strategies, Adapter and LoRA. The results show that FISGuard reduces the ProjRes attack AUC to near the random-guessing level of 0.5 in most settings, while maintaining downstream task performance close to that of the undefended model and introducing only limited computational overhead, thereby achieving a favorable privacy--utility trade-off.

↑