概率AI模型的TP-CRIV中统计可分性的表征
Characterizing Statistical Separability in TP-CRIV for Probabilistic AI Models
- Tohoku University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对概率AI模型的TP-CRIV,表征其统计可分性,关联匹配与非匹配证明者的逐挑战行为和验证级可分性,通过LLM实验验证了理论与经验AUC的一致性,为第三方验证提供统计基础。
AI中文摘要:
第三方挑战-响应身份验证(TP-CRIV)允许独立验证者在不直接访问参考模型的情况下,评估声明者是否拥有与远程部署模型完全相同的模型。然而,对于概率AI模型,同一查询的重复执行可能产生不同的输出,进而产生不同的验证观测结果。这引发了一个问题:应如何累积此类随机证据,以及需要多少证据才能实现可靠验证。在本研究中,我们对概率AI模型的TP-CRIV中的统计可分性进行了表征。具体而言,我们将匹配证明者与非匹配证明者的逐挑战行为与验证级可分性关联起来。该表征明确描述了独立挑战的数量和重复响应如何影响检测性能,并能够估算达到目标AUC所需的验证预算。我们使用开放式挑战为大语言模型(LLM)实例化了所提出的表征。实验结果显示存在匹配-非匹配分离,理论AUC与经验AUC高度吻合,且最小验证预算的估算结果一致。这些结果为将概率模型行为与验证级可分性及第三方验证所需证据关联起来提供了统计基础。
英文摘要:
Third-party challenge-response identity verification (TP-CRIV) enables an independent verifier to assess whether a claimant possesses a model identical to a remotely deployed model without directly accessing the reference model. However, for probabilistic AI models, repeated executions of the same query may produce different outputs and therefore different verification observations. This raises the question of how such stochastic evidence should be accumulated and how much evidence is required for reliable verification. In this work, we characterize statistical separability in TP-CRIV of probabilistic AI models. Specifically, we relate challenge-wise behavior of matching and non-matching provers to verification-level separability. The characterization explicitly describes how the numbers of independent challenges and repeated responses affect detection performance and enables the verification budget required for a target AUC to be estimated. We instantiate the proposed characterization for LLMs using open-ended challenges. The experiments demonstrate matching-non-matching separation, close agreement between theoretical and empirical AUCs, and consistent estimates of the minimum verification budgets. These results provide a statistical basis for relating probabilistic model behavior to verification-level separability and the evidence required for third-party verification.