arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越问答匹配:基于概率分布的大语言模型扰动-响应指纹识别

Beyond QA Matching: Perturbation-Response Fingerprinting via Probability Distributions for Large Language Models

Jichao Zeng, Yanli Chen, Hanzhou Wu

arXiv 2609.06330首次发表:更新:

发表机构

Guizhou Normal University; Shanghai University(贵州师范大学; 上海大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出BReF,一种无需训练的指纹方法,通过比较概率分布在扰动下的响应方向,在34个检查点基准上以MRR=1.0000准确识别模型来源,优于仅基于幅度的对照。

AI 中文摘要

大型语言模型经常经过指令微调、专业化、量化或其他变换,使得细粒度的来源追踪变得困难。在本文中,我们提出了BReF,一种无需训练的指纹识别方法,它比较在受控文本扰动下四个答案选项标签A/B/C/D上的概率分布如何移动。对于每对模型,BReF选择25个联合响应的探针,并通过全局余弦相似度比较它们的扰动对数比率(PLR)响应方向。在一个包含34个检查点、22个记录的直接父代关系和411个嫌疑-候选对的统一基准上,BReF在22/22的情况下检索到记录的父代(MRR=1.0000),DP-DF AUC为1.0000。同族判别更具挑战性(DP-SF AUC为0.8969),配对测试显示,与仅基于幅度的Top-25对照相比,在精确检索方面有显著提升。结合静态、随机探针、置换、校准和变换级别的对照,结果表明,强池化分离并不能保证在密切相关的检查点之间正确排序父代,验证了我们工作的优越性。

英文摘要

Large language models are often instruction-tuned, specialized, quantized, or otherwise transformed, making fine-grained provenance difficult. In this paper, we introduce BReF, a training-free fingerprint that compares how probability distributions over four answer-option labels A/B/C/D move under controlled textual perturbations. For each pair of models, BReF selects 25 jointly responsive probes and compares their perturbation log-ratio (PLR) response directions by global cosine similarity. On a unified benchmark with 34 checkpoints, 22 documented direct-parent relations, and 411 suspect-candidate pairs, BReF retrieves the documented parent in 22/22 cases (MRR=1.0000), with DP-DF AUC 1.0000. Same-family discrimination is harder (DP-SF AUC 0.8969), and paired tests show a significant exact-retrieval gain over a magnitude-only Top-25 control. Together with static, random-probe, permutation, calibration, and transformation-level controls, the results show that strong pooled separation does not guarantee correct parent ranking among closely related checkpoints, verifying the superiority of our work.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑