发表机构
University of California, Riverside(加州大学河滨分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出MineGrad攻击,可利用中毒预训练模型和微调参数从LoRA微调的共享梯度中分析性恢复用户数据,适用于语言和视觉任务,在多基准上实现高保真数据恢复,揭示了联邦微调的关键隐私漏洞。
AI 中文摘要
参数高效微调(PEFT),例如低秩适配(LoRA),最近已被应用于联邦学习以降低通信和计算成本。在该设置中,用户从服务器下载预训练模型后进行微调,在保持预训练模型冻结的同时,本地微调轻量级LoRA模块,仅将微调参数的梯度共享给服务器。尽管其日益流行,但针对对抗性服务器的联邦微调的鲁棒性仍未得到充分探索,服务器会恶意篡改训练协议以侵犯用户数据隐私。在本研究中,我们调查针对LoRA微调的梯度反演攻击,提出一种分析攻击,使恶意服务器能够利用中毒的预训练模型和微调参数恢复用户的私人数据。我们的设计将微调数据嵌入共享梯度中,以允许服务器分析性地重建用户数据。与先前的工作不同,我们的攻击适用于语言和视觉任务,不依赖于使用公共数据集进行计算成本高昂的(对抗性)预训练,也不要求训练令牌数量小于LoRA模块的秩。语言和视觉任务上的实验结果显示,在多个基准上实现了高保真数据恢复,揭示了几个关键漏洞。
英文摘要
Parameter-efficient fine-tuning (PEFT), such as low-rank adaptation (LoRA), has recently been adopted in federated learning to reduce communication and computation costs. In this setup, users download a pretrained model from the server prior to fine-tuning, and then fine-tune lightweight LoRA modules locally while keeping the pretrained model frozen, sharing only the gradients of the fine-tuning parameters with the server. Despite its growing popularity, robustness of federated fine-tuning against an adversarial server remains underexplored, where the server maliciously tampers with the training protocol to breach the privacy of users' data. In this work, we investigate gradient inversion attacks on LoRA fine-tuning. We propose an analytical attack that enables a malicious server to recover private user data by leveraging a poisoned pretrained model and fine-tuning parameters. Our design embeds fine-tuning data within the shared gradients, to allow the server to analytically reconstruct user data. Unlike prior works, our attack is applicable to both language and vision tasks, does not rely on computationally expensive (adversarial) pretraining with public datasets or require the number of training tokens to be less than the rank of LoRA modules. Experimental results on both language and vision tasks demonstrate high-fidelity data recovery across multiple baselines, revealing several critical vulnerabilities.
Comments2026 Annual Conference on Artificial Intelligence and Statistics (AISTATS 2026)