如何验证概率断言的一致性
How to Verify Probabilistic Consistency of Predictive Models
浏览论文内容
中文总结 AI 辅助
该研究针对概率预测器答案的自洽性验证问题,基于Nilsson的工作构造交互式PCP,为概率预测器自洽性证明提供复杂性理论基础,是训练模型证明自身一致性的第一步。
中文摘要 AI 辅助
当一个概率预测器回答大量条件概率查询时,其答案是否自洽,且能否在多项式时间内验证?该问题对AI安全具有重要意义,AI安全源于对AI行动可能导致的不良结果的概率预测保持诚实。我们按如下方式构造交互式PCP:设预测模型由概率电路P和输出预测置信度的电路Q指定,P和Q共同隐含指定了指数级多的概率断言。我们展示了一种协议,其中多项式时间验证器可验证(P,Q)的近似一致性。验证器获得电路对(P,Q),仅在少数点对其求值;同时获得证明预言机,即据称与(P,Q)预测一致的见证概率分布的编码,验证器在与单个不可信证明者交互时仅读取其少数位置。过程中,我们必须确保存在与模型预测一致的稀疏见证分布。为此,我们首先考虑显式概率断言一致性的见证分布,而非预测器指定的断言:例如在n个布尔变量上的m个断言,每个形式为Pr[Y=1 | X=x] = p。基于Nilsson(Artif. Intell., 1986)开创的工作,我们将显式断言的l₂近似概率一致性归入NP类,其证书长度为输入比特精度B的O(mn + log B);我们进一步证明,小的加性完备性-可靠性间隙可消除对B的依赖。这些结果共同为证明概率预测器的自洽性提供了复杂性理论基础。我们将交互式PCP视为训练预测模型以证明自身一致性的第一步。
英文摘要
When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of interest for AI safety, where safety is derived from honesty about probabilistic predictions of unwanted outcomes potentially caused by an AI action. We construct an interactive PCP as follows. Let a predictive model be specified by a probability circuit P and a circuit Q which outputs confidence in predictions. Together, P and Q implicitly specify exponentially many probabilistic claims. We show a protocol in which a polynomial-time verifier can verify the approximate consistency of (P,Q). The verifier is given the pair of circuits (P,Q), which it evaluates at only a few points; alongside them it is given a proof oracle, an encoding of a witnessing probability distribution allegedly consistent with the predictions of (P,Q), which it reads at a few locations while interacting with a single untrusted prover. En route, we must ensure the existence of a sparse witnessing distribution consistent with the model's predictions. To do so, we first consider witness distributions for the consistency of explicit probabilistic claims, rather than claims specified by a predictor: say m claims, each of the form Pr[Y = 1 | X = x] = p, over n Boolean variables. Building on work initiated by Nilsson (Artif. Intell., 1986), we place l_2-approximate probabilistic consistency of explicit claims in NP, with certificates of length O(mn + log B) in the input bit-precision B; we further show how a small additive completeness-soundness gap removes the dependence on B. Together these results provide a complexity-theoretic foundation for certifying the self-consistency of probabilistic predictors. We view our interactive PCP as a first step toward training predictive models to prove their own consistency.
发表机构
- EPFL(洛桑联邦理工学院)
- Université de Montréal(蒙特利尔大学)
- Mila -- Quebec AI Institute(米拉-魁北克人工智能研究所)
- LawZero
- MIT(麻省理工学院)
- UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。