梯度对分割语言模型中文本泄露的贡献:按词元与按文档计数
What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document
浏览论文内容
中文总结 AI 辅助
本研究量化了分割学习中梯度对文本泄露的额外贡献,发现按文档计数时泄露显著增加,并建议同时按词元和文档报告泄露,将分割模型发送的数据视为与文本同等敏感。
中文摘要 AI 辅助
分割学习使客户端能够在服务器上训练语言模型,而无需发送其文本。客户端自行运行前若干层,仅将每层的输出(每个词元对应的数值向量)发送给服务器。在训练过程中,服务器会回传梯度。我们证明,分割点处的观察者可以从这些流量中重建客户端的大部分文本,并量化了梯度对此的贡献程度。在GPT-2上,仅持有客户端各层公开权重(publicly released weights)的攻击者,仅凭激活值即可恢复94.20%的词元;若同时看到梯度,恢复率可达97.38%,提升了3.17个百分点(95%置信区间为[2.72, 3.64])。按文档计数时,差异更为显著。攻击者在不借助梯度的情况下,能精确重建13.71%的32词元文档;而借助梯度时,该比例升至37.77%,因为只有当每个词元都正确时,文档才算重建成功。计数方式也会影响防御效果的评价。秘密混合(Secret mixup)方法将每个输出向量与诱饵混合,能阻止攻击者精确重建几乎任何文档,但攻击者仍能恢复83-91%的词元。在第二个实验中,使用GPT-2和Qwen3-0.6B,服务器仅训练一段连续的层,该段起始层的位置会同时影响模型质量和泄露程度,即使该段长度固定也是如此。我们建议同时按词元和按文档报告泄露情况,并将分割模型发送的数据视为与文本本身同等敏感。
英文摘要
Split learning lets a client train a language model on a server without sending its text. The client runs the first layers itself and sends the server only their output, a vector of numbers for each token. During training, the server sends gradients back. We show that an observer at the split can rebuild most of the client's text from this traffic, and we measure how much the gradients help. On GPT-2, an attacker who holds only the publicly released weights of the client's layers recovers 94.20% of tokens from the activations alone and 97.38% when it also sees the gradients, 3.17 percentage points more 95% interval [2.72, 3.64]. Counted by document, the difference is much larger. The attacker rebuilds 13.71% of 32-token documents exactly without the gradients and 37.77% with them, because a document only counts when every token is right. How we count also changes how good a defence looks. Secret mixup, which blends each outgoing vector with a decoy, stops the attacker from rebuilding almost any document exactly, yet the attacker still recovers 83-91% of tokens. In a second experiment on GPT-2 and Qwen3-0.6B, where the server trains only a run of consecutive layers, the layer at which the run starts changes both model quality and leakage, even when the run's length is fixed. We recommend reporting leakage both per token and per document, and treating what a split model sends as being as sensitive as the text itself.