Witeness Overlap:开放权重模型家族内部的定向来源追踪
Witeness Overlap: Directional Provenance Inside Open-Weight Model Families
AI总结:
针对开放权重模型来源审计中方向不明的问题,提出Witness Overlap方法,利用同族第三方检查点进行局部几何比较,以无提示、无训练的白盒方式定向父子关系,在176个检查点上达到95.3%的准确率。
AI中文摘要:
开放权重模型经常被发布、微调、对齐、合并和重新发布,这使得来源审计不仅要询问检查点之间是否相关,还要确定哪个检查点更早出现。许多现有的模型来源追踪方法是为已知基础模型的审计场景设计的:给定一个受害者或源模型,它们测试嫌疑模型是否与之相关。尽管这些审计被框架化为从源到嫌疑的测试,但其底层证据往往是对称的,依赖于表示相似性、权重相似性、行为指纹或相关性统计。对称的成对比较可以检测相关性,但无法单独确定检查点A和B之间的关系方向。因此,我们引入了一种局部几何比较方法:不是直接比较两个检查点,而是添加同一家族的第三个检查点作为见证者,并比较每个候选端点周围的几何结构。通过询问哪个候选者更像分支父节点来推断方向。受此想法以及观察到的父锚定与子锚定的见证重叠分布之间的经验不对称性的启发,我们提出了Witness Overlap,一种无需提示、无需训练的白盒测试,用于定向来源追踪。在来自16个家族的176个LLM检查点上,我们的单见证测试使用Frobenius余弦定向了95.3%的父子决策。我们进一步评估了根识别、兄弟区分、对VLM和扩散家族的泛化以及链式结构排序。该信号对权重噪声和稀疏剪枝具有鲁棒性,提出的SVD权重降维变体比Frobenius余弦表现出更强的鲁棒性。
英文摘要:
Open-weight models are often released, fine-tuned, aligned, merged, and re-released, making provenance audits ask not only whether checkpoints are related, but also which checkpoint came first. Many existing model-provenance methods are designed for a base-known audit setting: given a victim or source model, they test whether a suspect model is related to it. Although these audits are framed as source-to-suspect tests, their underlying evidence is often symmetric, relying on representation similarity, weight similarity, behavioral fingerprints, or correlation statistics. Symmetric pairwise comparisons can detect relatedness, but they cannot by themselves orient relationship between checkpoints A and B. We therefore introduce a local geometric comparison: instead of comparing two checkpoints directly, we add a third same-family checkpoint as a witness and compare the geometry around each candidate endpoint. Direction is inferred by asking which candidate behaves more like a branching parent. Motivated by this idea, and by the empirically observed asymmetry between parent-anchored and child-anchored witness-overlap distributions, we propose Witness Overlap, a prompt-free, training-free white-box test for directional provenance. On 176 LLM checkpoints from 16 families, our one-witness test orients 95.3\% of parent-child decisions using Frobenius cosine. We further evaluate root identification, sibling discrimination, generalizations to VLM and diffusion families, and chain-structured ordering. The signal is robust to weight noise and sparse pruning, with a proposed SVD weight reduction variant showing greater robustness than Frobenius cosine.