arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于归一化响应相似性的回顾性蒸馏归因

Retrospective Distillation Attribution via Normalized Response Similarity

Minwoo Jang, Jaechang Kim, Minhyeon Oh, Jeongyeon Hwang, Jungseul Ok

arXiv 2609.32749首次发表:更新:

AI 中文总结

提出仅基于输出的SCOUT方法,通过聚合句法模式并校准距离,实现对经历后续训练的蒸馏模型进行有效溯源归因,并发现教师句法特征在训练中持续存在。

AI 中文摘要

模型蒸馏通过在教师响应上进行监督微调(SFT)来迁移能力,这些响应通常从商业API收集,从而引发了模型溯源问题。现有的蒸馏归因方法大多在SFT步骤后立即对学生模型进行评估。然而,蒸馏模型在发布前可能经历进一步的SFT、偏好优化或强化学习,而审计者可能无法访问基于参考的归因所需的预蒸馏检查点。为弥补这一差距,我们提出了SCOUT,一种仅基于输出的方法,该方法将重复出现的句法模式聚合为候选档案,过滤低对比度模式,并根据候选间距离校准学生-候选距离。SCOUT仅使用当前文本即可支持归因和弃权(不执行),无需模型权重、词元似然或历史检查点。对涵盖多种不同训练后目标的蒸馏模型公开发布后代的审计中,SCOUT始终能识别蒸馏源。此外,沿着训练轨迹追踪与教师相关的句法特征表明,这些特征在蒸馏期间出现,并在随后的偏好优化和强化学习中持续存在。

英文摘要

Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However, a distilled model may undergo further SFT, preference optimization, or reinforcement learning before release, while an auditor may lack access to the pre-distillation checkpoint required by reference-based attribution. To close this gap, we propose SCOUT, an output-only method that aggregates recurring *syntactic patterns* into candidate profiles, filters low-contrast patterns, and calibrates student--candidate distances against inter-candidate distances. SCOUT supports attribution and abstention using only current texts, without model weights, token likelihoods, or historical checkpoints. Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source. Furthermore, tracing teacher-associated *syntactic signatures* along training trajectories reveals that they emerge during distillation and persist through subsequent preference optimization and reinforcement learning.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑