arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型中基于参考的蒸馏检测

Reference-Based Distillation Detection in LLMs

Rajat Rawat, Sizhe Chen, Akshay Anand, Michael Duan, Bob Rotsted, Sewon Min

arXiv 2607.09692首次发表:更新:

发表机构

University of California, Berkeley; University of Southern California; OpenAI(加利福尼亚大学伯克利分校; 南加利福尼亚大学; OpenAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在检测大语言模型是否经蒸馏而来,提出基于参考的成员推理蒸馏检测方法,通过比较模型与不同候选教师输出的优先对齐程度识别教师及蒸馏证据,能处理未知管道,扩展到开放世界设置,应用于当代模型获潜在蒸馏关系新证据。

AI 中文摘要

模型蒸馏(在更强的第三方模型输出上进行训练)被广泛用于提升性能,但引发了对不公平优势和违反政策的担忧。这促使一个基本问题:能否检测一个模型是否从另一个模型蒸馏而来?研究表明,孤立地从学生模型识别教师模型极具挑战,但在基于参考的设置中变得可行。给定一个模型和同一谱系的早期检查点,能识别用于训练后期检查点的教师模型。介绍了一种基于参考的成员推理的蒸馏检测方法,通过比较学生模型相对于参考检查点与不同候选教师输出的优先对齐程度来识别最可能的教师并检测蒸馏证据。为处理未知蒸馏管道,直接从模型输出推断代理提示模板,还识别了o1/o3模型特有的字形级信号。由于现代模型谱系已严重纠缠,评估蒸馏检测具有挑战性,为此开发了涵盖受控蒸馏实验和真实世界模型的混合评估。在两种设置下,该方法在单教师蒸馏场景中以近乎完美的准确率恢复真实教师,即便基础蒸馏管道很大程度未知。还引入了教师归因和蒸馏检测的统计测试,并将框架扩展到开放世界设置。将方法应用于当代模型,得到了关于QwQ、DeepSeek-R1和GPT-OSS潜在蒸馏关系的新证据。

英文摘要

Model distillation -- training on outputs from stronger third-party models -- is widely used to boost performance, but raises concerns about unfair advantages and policy violations. This motivates a fundamental question: can we detect whether a model was distilled from another? We show that, while identifying a teacher model from a student in isolation is highly challenging, it becomes tractable in a reference-based setting: given a model and an earlier-generation checkpoint from the same lineage, we can identify the teacher model used to train the later checkpoint. We introduce a distillation detection method based on reference-based membership inference. By comparing how strongly a student model preferentially aligns with outputs from different candidate teachers relative to a reference checkpoint, our method identifies the most likely teacher and detects evidence of distillation. To handle unknown distillation pipelines such as hidden prompts, we infer proxy prompt templates directly from model outputs. We additionally identify a distinctive glyph-level signal specific to o1/o3 models. Evaluating distillation detection is challenging because modern model lineages are already heavily entangled. To address this, we develop a hybrid evaluation spanning both controlled distillation experiments and real-world models. Across both settings, our approach recovers the true teacher with near-perfect accuracy in single-teacher distillation scenarios, even when the underlying distillation pipeline is largely unknown. We further introduce statistical tests for both teacher attribution and distillation detection, and extend our framework to open-world settings where no teacher is guaranteed to be present among the candidates. Applying our method to contemporary models yields new evidence regarding potential distillation relationships involving QwQ, DeepSeek-R1, and GPT-OSS.

Comments27 pages (14 main), 21 figures, 16 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑