arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04382cs.CRcs.DCcs.LG

Split-LLM训练中的隐私漏洞:返回的梯度抵消了诱饵

Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys

Georgios Politis, Evangelos Pappas

首次发表
浏览论文内容

中文总结 AI 辅助

本研究发现Split-LLM训练系统的返回梯度存在隐私漏洞,零值模式可暴露真实行,裁剪加噪仅能部分缓解,且存在未检测的攻击类别。

中文摘要 AI 辅助

我们对两节点Split-LLM训练系统开展了系统安全案例研究,该系统的隐私评估通过了测试,但存在一条可观测信道未被检测。可信本地节点(TLN)向不可信云节点(UCN)发送受保护的激活值,UCN返回其输出,持有私有损失的TLN返回输出梯度。UCN接收的帧混合了真实行与诱饵,而损失会忽略诱饵,诱饵对应的梯度恰好为零,因此零值模式会暴露哪些行是真实行。我们采用预先固定的协议进行测量:注入已知强度的泄露以证明仪器可检测到泄露,使用打乱标签的对照组以证明仪器不会报告不存在的泄露,且在运行前设定阈值。在9个随机种子下,零值在每个帧中都识别出了真实行,每次运行中4096行里的4096行均为真实行。针对帧内容的攻击相比常数猜测基准多恢复了约每百个令牌1个额外令牌(提升0.65至1.50个百分点);打乱标签的对照组未恢复任何内容。第二组运行在保持模型质量在预算范围内的配置上重复了该实验,因此该发现并非局限于无人会部署的设置。在两个数据集上,此类运行均通过了前向信道隐私检查和质量检查,但当包含返回梯度后,该检查失败。对梯度的每一行进行裁剪和加噪可关闭泄露,仅需约0.01 nat的保留交叉熵损失。但该系统并未因此安全:五类攻击(包括那些跨训练步骤累积观测值的攻击)从未被测量。

英文摘要

We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested. The Trusted Local Node (TLN) sends protected activations to the Untrusted Cloud Node (UCN), the UCN returns its output, and TLN, holding the private loss, returns the output gradient. The frame the UCN receives mixes real rows with decoys, and the loss ignores the decoys. Their gradients are exactly zero, so the pattern of zeros reveals which rows were real. We measure it with a protocol fixed in advance: a leak injected at known strength to prove the instrument can see one, a shuffled-label control to prove it does not report absent leaks, and a threshold set before the runs. Across nine seeds, the zeros identified the real rows on every frame, 4,096 of 4,096 per run. An attack on the frame contents recovered about one extra token per hundred over a constant-guess baseline (+0.65 to +1.50 percentage points); the shuffled controls recovered nothing. A second set of runs repeated this on a configuration that keeps model quality within budget, so the finding is not confined to a setting nobody would deploy. On both datasets, every such run passed the forward-channel privacy check and the quality check, yet failed that same check once the returned gradient was included. Clipping and noising each row of the gradient closed the leak for about 0.01 nats of held-out cross-entropy. The system is not thereby safe: five classes of attack, including those accumulating observations across training steps, were never measured.

↑