arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Stemma:诱导决策区域揭示大语言模型的来源

Stemma: Induced Decision Regions Reveal LLM Provenance

Keyu Zhang, Vadim Safronov, Andrew Martin

arXiv 2607.25880首次发表:更新:

发表机构

University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型来源测试中现有黑盒方法可靠性不足的问题,引入诱导决策区域,提出Stemma方法,以稳定性等为探针选择原则,在多组源-可疑对测试中表现出色,显著优于基线,对不同部署设置具鲁棒性。

AI 中文摘要

大语言模型来源测试旨在判断可疑大语言模型是否与源模型属于同一系列。现有的黑盒方法大多从响应层面特征推断这种关系,但这些特征在适应或部署时可能变化,削弱了来源证据的可靠性。为解决此局限,我们通过将开放式输出映射到有限决策空间引入诱导决策区域,将来源测试重新定义为测量决策区域的继承性。实证分析表明相关模型中源模型的诱导区域保留更强。在此基础上,我们提出Stemma,一种实用的黑盒大语言模型指纹识别方法,以稳定性、鲁棒性和特异性为互补探针选择原则,可靠估计诱导决策区域继承性。在770个源-可疑对及1260个对的测试中,Stemma在1%误报率下分别达到0.967的AUC和87.8%的真阳性率,以及0.995的AUC和93.5%的真阳性率,显著优于四个代表性基线,对不同推理时部署设置具有鲁棒性。

英文摘要

LLM provenance testing asks whether a suspect LLM belongs to the same lineage as a source. Existing black-box methods largely infer this relationship from response-level characteristics, but these characteristics may shift under adaptation or deployment even when the underlying meaning remains unchanged, weakening the reliability of provenance evidence. To address this limitation, we introduce induced decision regions by mapping open-ended outputs into a finite decision space, thereby abstracting away surface-form variation and reframing provenance testing as measuring the inheritance of decision regions. Empirical analysis shows that the source's induced regions are preserved more strongly in related models than in unrelated models. Building on this signal, we propose Stemma, a practical black-box LLM fingerprinting method that operationalises stability, robustness, and specificity as complementary probe-selection principles for reliably estimating induced decision region inheritance. Across 770 source-suspect pairs drawn from 56 public checkpoints and spanning diverse model-weight transformations, Stemma achieves 0.967 AUC and 87.8% TPR at 1% FPR, substantially outperforming four representative baselines. It further achieves 0.995 AUC and 93.5% TPR at 1% FPR on 1,260 pairs covering 91 deployment instances, demonstrating robustness to diverse inference-time deployment settings.

Comments25 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑