arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

海报:大语言模型蒸馏推理的初步研究

Poster: A Preliminary Study of LLM Distillation Inference

Edward Chen, Yuntao Du

arXiv 2610.12137首次发表:更新:

发表机构

Carmel High School; Purdue University(卡梅尔高中; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对未经授权的LLM蒸馏威胁,提出通过假设检验结合影子模型的蒸馏推理方法,在Qwen2.5-7B与Llama-3.2-3B的实验中实现了0.02显著性水平下1.0的真阳性率,验证了该检测方法的可行性。

AI 中文摘要

未经授权的模型蒸馏是指在专有大语言模型(LLM)的输出上训练模型,这对模型提供者而言是日益严重的威胁。我们研究蒸馏推理:判断可疑模型是由另一模型蒸馏而来还是独立训练的。我们将此问题表述为假设检验,并通过训练影子模型估计每个假设下的预期行为:蒸馏影子模型从教师模型的推理轨迹中学习,而独立影子模型仅从参考答案中学习。审计员测量可疑模型预测教师模型推理输出的接近程度,随后利用影子模型将可疑模型的分数转换为校准后的p值。在以Qwen2.5-7B为教师模型、Llama-3.2-3B为可疑模型的初步研究中,我们的测试在显著性水平0.02下达到了1.0的真阳性率。这些结果证明了利用蒸馏推理检测蒸馏攻击的可行性。

英文摘要

Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers. We study distillation inference: determining whether a suspect model was distilled from another model or trained independently. We formulate this problem as a hypothesis test and estimate the behavior expected under each hypothesis by training shadow models: distilled shadow models learn from the teacher's reasoning traces, whereas independent shadow models learn only from reference answers. The auditor measures how closely each model predicts the teacher's reasoning outputs and then uses the shadow models to convert the suspect's score into a calibrated p-value. In a preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, our test achieves a true positive rate of 1.0 at a significance level of 0.02. These results demonstrate the feasibility of using distillation inference to detect distillation attacks.

CommentsAccepted as a poster paper at the 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS'26)

DOI:10.1145/3830454.3846416

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑