arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当路由暴露成员身份:来自MoE路由器遥测的隐私泄露

When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry

Yixin Tan, Jiayang Liu, Lu Sun, Yuke Hu, Zheng Li, Rui Wen

arXiv 2610.10616首次发表:更新:

发表机构

Institute of Science Tokyo; Nanyang Technological University; Tohoku University; King Abdullah University of Science and Technology; Shandong University(东京科学大学; 南洋理工大学; 东北大学; 阿卜杜拉国王科技大学; 山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出路由器增强的成员推理攻击,发现MoE模型的路由器遥测可泄露样本是否用于微调,在多设置下提升成员推理性能,是微调MoE模型的额外隐私风险面。

AI 中文摘要

混合专家(MoE)语言模型在推理过程中会产生路由信息,这些信息可能被记录或暴露,用于监控、调试、负载分析和安全审计。与普通模型输出不同,这种遥测信息揭示了模型内部计算的视图,引发了一个隐私问题:它能否揭示某个样本是否被用于微调部署后的模型?我们提出了一种路由器增强的成员推理攻击,该攻击将传统的输出侧信号与聚合的路由特征相结合,并将从独立微调的影子模型中学到的成员分类器应用于目标模型。在三种MoE架构和三个数据域中,路由器遥测始终能比强大的输出信号集成方法提升成员推理性能,在所有九种设置下,当FPR为1%时,TPR提升了2.7至9.4个百分点。这种泄露在全微调、路由器冻结训练、LoRA和指令微调中均存在,且仅通过离散的专家选择、受限的遥测或单个影子模型即可观测到。机制分析进一步表明,这种泄露不需要路由器特定的记忆:微调会将成员信息引入隐藏表示,而即使路由器参数冻结,路由器也会暴露该信号的投影。对遥测进行扰动会随着其保真度的降低而减少这种额外泄露。我们的结果表明,路由器遥测可将操作信号转化为微调后MoE模型的额外隐私面。

英文摘要

Mixture-of-Experts (MoE) language models produce routing information during inference that may be logged or exposed for monitoring, debugging, load analysis, and safety auditing. Unlike ordinary model outputs, this telemetry reveals a view of the model's internal computation, raising a privacy question: can it reveal whether an example was used to fine-tune the deployed model? We introduce a router-augmented membership inference attack that combines conventional output-side signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to the target model. Across three MoE architectures and three data domains, router telemetry consistently improves membership inference over a strong output-signal ensemble, increasing TPR at 1\% FPR by 2.7--9.4 percentage points across all nine settings. The leakage persists across full fine-tuning, frozen-router training, LoRA, and instruction tuning, and remains observable with only discrete expert selections, restricted telemetry, or a single shadow model. Mechanistic analysis further shows that the leakage does not require router-specific memorization: fine-tuning introduces membership information into hidden representations, while the router exposes a projection of this signal even when its parameters are frozen. Perturbing the telemetry reduces this additional leakage only as its fidelity degrades. Our results show that router telemetry can turn an operational signal into an additional privacy surface for fine-tuned MoE models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑