发表机构
Amazon; Massachusetts Institute of Technology(亚马逊; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出中心化残差特征方法,通过移除残差训练产生的共享身份对齐组件,对比残差块特有结构得到谱系分数,可高效准确验证开源权重语言模型检查点的谱系
AI 中文摘要
开源权重语言模型会经过微调、量化、剪枝和合并,但其谱系往往未被记录。我们研究无数据白盒谱系验证:仅通过权重能否判断两个兼容模型检查点是否具有共同谱系?残差训练会在分支产物中产生共享的身份对齐组件,仅靠该结构无法确定谱系。我们移除该组件,对比残差块中检查点特有的结构,得到针对独立检查点校准的对称谱系分数。在残差MLP和GPT-2基准上,该分数能将微调、LoRA合并、剪枝和量化的后代与独立及蒸馏模型区分开(AUROC=1.0),将权重谱系与行为相似性区分开。在保留功能的检查点清洗实验中,权重空间基线会丧失优势或失效;我们的分数保持不变,且在GPT-2上比最近的鲁棒基线快76倍。该投影配对信号在六个语言模型家族及其他场景中均存在,案例研究正确识别了3个相关和7个不相关的LLaMA-2公开检查点。总体而言,这些结果为兼容的开源权重语言模型检查点建立了一种被动、无数据的谱系信号
英文摘要
Open-weight language models are fine-tuned, quantized, pruned, and merged, yet their provenance is often undocumented. We study data-free white-box lineage verification: can weights alone reveal whether two compatible model checkpoints share ancestry? Residual training produces a shared identity-aligned component in branch products, so this structure alone cannot establish ancestry. We remove it and compare checkpoint-specific structure across residual blocks, yielding a symmetric lineage score calibrated against independent checkpoints. On residual-MLP and GPT-2 benchmarks, the score separates fine-tuned, LoRA-merged, pruned, and quantized descendants from independent and distilled models (AUROC=1.0), distinguishing weight ancestry from behavioral similarity. Under function-preserving checkpoint laundering experiments, weight-space baselines lose margin or fail; our score remains unchanged and runs 76x faster than the nearest robust baseline on GPT-2. The projection-pairing signal appears across six language-model families and beyond, and a case study correctly identifies 3 related and 7 unrelated LLaMA-2 public checkpoints. Collectively, these results establish a passive, data-free provenance signal for compatible open-weight language-model checkpoints
CommentsPreprint