AI 中文总结
研究开放权重语言模型谱系验证问题,提出modelDNA工具,通过HTTP读取指纹识别模型,与参考数据库比较并返回判定类别,实现高AUROC和零误报,还能进行合并分解,恢复混合权重,公开相关数据可离线重现。
AI 中文摘要
开放权重语言模型的谱系图是自我报告的:Hugging Face的base_model元数据字段是可选且未经验证的,超过60%的Hub模型根本没有记录其来源。研究文献中存在从权重检测谱系的方法,但每个方法都作为与一个信号和一个实验相关的论文代码发布;当出现来源争议时,分析需要手动重新进行。本报告介绍了modelDNA,一种通过大约100 - 300MB的范围HTTP读取(而不是为7B模型完整下载15GB)对模型进行指纹识别的工具,将指纹与四个已发布信号家族的基础模型参考数据库进行比较,并以校准概率返回八个判定类别之一,宁愿诚实弃权也不愿自信犯错。在15个有组织记录来源的真实Hub模型的基准测试中,与8个候选基础模型(13个阳性,107个硬阴性)进行比较,该系统实现了AUROC 1.0,在报告阈值处零误报,并且13/13的正确top - 1父系归属。报告的第二个贡献是合并分解。每个主流的权重合并方法在每个张量上都是(近)线性的,并且指纹样本位置是张量标识的确定性函数,因此合并模型的指纹是其父模型指纹的相同线性组合。因此,可以通过和为1约束的最小二乘法仅从指纹中恢复混合权重。与以发布的mergekit配置作为地面真值的合并进行比较,该方法在r = 0.999时恢复了slerp合并的层插值曲线,并且恢复了dare_ties合并的混合权重,与发布值的误差在0.011以内,无需下载指纹之外的任何权重。所有55个模型的指纹、基准测试和推断的谱系图都是公开的且可离线重现。
英文摘要
The lineage graph of open-weight language models is self-reported: Hugging Face's base_model metadata field is optional and unverified, and over 60% of Hub models document no parentage at all. Methods for detecting lineage from weights exist in the research literature, but each ships as paper code tied to one signal and one experiment; when a provenance dispute breaks, the analysis is redone by hand. This report describes modelDNA, a tool that fingerprints a model from roughly 100-300 MB of ranged HTTP reads (instead of a full 15 GB download for a 7B model), compares the fingerprint against a reference database of foundation models across four published signal families, and returns one of eight verdict classes with a calibrated probability, preferring honest abstention to confident error. On a benchmark of 15 real Hub models with org-documented parentage, judged against 8 candidate bases (13 positives, 107 hard negatives), the system achieves AUROC 1.0, zero false positives at its reporting threshold, and 13/13 correct top-1 parent attribution. The report's second contribution is merge decomposition. Every mainstream weight-merging method is (near-)linear per tensor, and fingerprint sample positions are deterministic functions of tensor identity, so a merged model's fingerprint is the same linear combination of its parents' fingerprints. Mixture weights can therefore be recovered from fingerprints alone by sum-to-one constrained least squares. Against merges with published mergekit configurations as ground truth, the method recovers a slerp merge's layer-interpolation curves at r = 0.999 and a dare_ties merge's mixture weights to within 0.011 of the published values, without downloading any weights beyond the fingerprints. All fingerprints, benchmarks, and the inferred lineage graph of 55 models are public and reproducible offline.
CommentsCode: https://github.com/AwaisAdilKhokhar/modelDNA . Data: https://huggingface.co/datasets/AwaisAdilKhokhar/modeldna-atlas . Live scanner: https://huggingface.co/spaces/AwaisAdilKhokhar/modelDNA . DOI: 10.5281/zenodo.21305586