发表机构
EPFL; ENS Paris-Saclay; Université Paris-Saclay; MATS(洛桑联邦理工学院; 巴黎萨克雷高等师范学校; 巴黎萨克雷大学; MATS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Diff Mining框架,通过对比微调模型与基础模型的对数来识别微调目标,在微调领域检测、偏见识别等任务上优于现有方法,可用于开发微调审计工具。
AI 中文摘要
微调已成为优化语言模型现有行为并诱导新行为的黄金标准,但在此过程中究竟会出现哪些行为往往仍不明确。随着模型能力不断提升,更好地理解微调变得愈发重要,尤其是因为微调过程中可能会出现不受欢迎的行为。本文中,我们提出Diff Mining这一简单却有效的框架,通过将微调模型的对数(logits)与其基础模型的对数进行比较,来识别微调模型学到的内容。Diff Mining能有效凸显微调模型中被放大的显著 token,即便在与微调领域无关的文本上,也可作为其训练的指纹。与许多需要访问模型内部结构的现有模型差异方法不同,Diff Mining仅需访问输出对数,且可扩展到大型模型。该框架包含两个模块化阶段:(i)在参考语料库上提取微调模型与基础模型之间每个上下文的对数差;(ii)聚合所得信号以构建代表微调过程的可解释 token 集合。对于聚合,我们探索了两种方法:简单的 Top-K 频率方法,以及基于非负矩阵分解(Non-negative Matrix Factorization,NMF)的方法,用于将多个微调目标解耦为不同的 token 簇。实证结果表明,Diff Mining在多种场景中均表现出色:在微调领域检测任务上,无论是识别相关 token,还是当可解释智能体获得提取的 token 集后进行下游性能测试,它都显著优于最先进的模型差异方法;在带有注入偏见的模型上,它无需针对性探测即可识别出超过三分之一的偏见。总体而言,我们的框架在开发用于检测微调目标的审计工具方面展现出良好前景。
英文摘要
Finetuning has become the gold standard for refining existing behaviors and inducing new ones in language models, yet it often remains unclear exactly which behaviors emerge during this process. As models grow ever more capable, understanding finetuning better becomes increasingly important, particularly since unwanted behaviors may arise during finetuning. In this paper, we introduce Diff Mining, a simple yet effective framework for identifying what a finetuned model has learned by comparing its logits to those of its base model. Diff Mining effectively surfaces salient tokens that are amplified in the finetuned model, serving as a fingerprint of its training -- even on text unrelated to the finetuning domain. Unlike many existing model diffing methods which require model internals, Diff Mining only needs access to output logits and scales to large models. The framework consists of two modular stages: (i) extracting per-context logit differences between the finetuned and base models on a reference corpus, and (ii) aggregating the resulting signals to construct an interpretable token set representing the finetune. For aggregation, we explore both a simple Top-K frequency method and a Non-negative Matrix Factorization (NMF)-based approach for disentangling multiple finetuning objectives into distinct token clusters. Empirically, Diff Mining succeeds across diverse settings: on finetune domain detection, it significantly outperforms state-of-the-art model diffing methods both in identifying relevant tokens and in downstream performance when an interpretability agent is given access to the extracted token set; on models with injected biases, it identifies more than one third of the biases without targeted probing. Overall, our framework shows promise in developing auditing tools to detect finetuning objectives.
Comments37 pages, 7 figures. ICLR 2026 Workshop: Principled Design for Trustworthy AI. Code available at https://github.com/science-of-finetuning/diffing-toolkit