AI 中文总结
针对大语言模型拆分学习中梯度匹配攻击的标签泄露问题,提出Gradient Mirage防御方法,通过三个维度的不一致性打破梯度-目标一致性,在保留优化效用的同时提升隐私保护,实现更优的隐私-效用权衡。
AI 中文摘要
大语言模型拆分学习(SL)中的梯度匹配攻击(GMAs)依赖于一个关键但未被充分探讨的假设:在拆分接口处暴露的梯度是客户端全标签训练目标的忠实导数。这种梯度-目标一致性使得好奇的服务器能够通过搜索能产生观测梯度的序列来恢复私有标签。我们提出了Gradient Mirage,一种不丢弃反向信号优化效用的防御方法,它打破了这种一致性。我们的核心思路是诱导攻击者求解一个错误设定的逆问题,使得序列空间中没有合理的标签序列能解释观测到的梯度。具体而言,Gradient Mirage通过在目标、方向和规模三个维度诱导不一致来实现这一点:选择性自回归监督从掩码代理损失中导出暴露梯度,而非攻击者假设的全标签目标;规模遮蔽则应用随机乘性重缩放,掩盖梯度的自然幅度;方向私有化进一步通过von Mises-Fisher(vMF)机制在方向度量差分隐私保证下随机化梯度方向,同时保持其幅度。关键是,效用得以保留:顶层仍通过双轨反向传播从所有目标令牌学习,暴露的梯度仍具有信息性,因为每个监督令牌都保留其完整的自回归上下文,底层梯度恢复则为底层分段优化恢复有效梯度。大量实验表明,在可比的微调性能下,Gradient Mirage提供了比现有防御强得多的保护,实现了更好的隐私-效用权衡。
英文摘要
Gradient matching attacks (GMAs) in LLM split learning (SL) rely on a critical yet underexplored assumption: the gradient exposed at the split interface is a faithful derivative of the client's full-label training objective. This gradient-objective consistency allows a curious server to recover private labels by searching for a sequence whose induced gradient explains the observation. We propose Gradient Mirage, a defense that breaks this consistency without discarding the optimization utility of the backward signal. Our key idea is to induce the adversary to solve a misspecified inverse problem, in which no plausible label sequence in the sequence space can explain the observed gradients. Concretely, Gradient Mirage achieves this by inducing inconsistency across three dimensions: objective, direction, and scale. Selective Autoregressive Supervision derives the exposed gradient from a masked surrogate loss rather than the full-label objective assumed by the attacker; Scale Blinding then applies randomized multiplicative rescaling, obscuring the gradient's natural magnitude; and Directional Privatization further randomizes the gradient direction while preserving its magnitude through the von Mises-Fisher (vMF) mechanism under a directional metric differential privacy guarantee. Crucially, utility is preserved: the Top segment still learns from all target tokens via Dual-Track Backpropagation, the exposed gradient remains informative since each supervised token retains its complete autoregressive context, and Bottom-Gradient Recovery restores the effective gradient for Bottom-segment optimization. Extensive experiments show that Gradient Mirage provides substantially stronger protection than existing defenses under comparable fine-tuning performance, achieving a better privacy-utility trade-off.