AI 中文总结
本研究探讨AI监督的零知识可能性,证明在一般预言机辅助计算中零知识论证不可实现,但若预言机附加签名,则可实现高效零知识验证,为可扩展监督提供新途径。
AI 中文摘要
AI 系统日益从机密数据中产生输出,例如根据医疗记录进行的工作适应性评估,或根据药物的秘密结构预测其候选药物的性质。在不泄露底层数据的情况下验证此类输出的正确性至关重要。最近的一系列工作研究了通过交互式证明和辩论对预言机辅助计算进行 AI 输出验证,其中正确性可能依赖于人类判断、物理实验或网络等预言机。这些工作侧重于由运行速度远快于计算的验证者进行验证。然而,对于一般的预言机辅助计算,这种高效验证是不可能的,因此这些工作依赖于额外的假设。我们转而关注隐私性:允许验证者在计算的多项式时间内运行,我们询问针对预言机辅助计算的交互式论证是否可以是零知识的,从而使验证者除了输出的正确性之外,对机密数据一无所知。我们证明,一般来说,这是不可能的。在随机预言机模型中,不存在针对所有预言机辅助计算的零知识证明,即使允许证明者和验证者运行的时间都比计算本身长得多。这种不可能性扩展到辩论,这是可扩展监督的典型模型。在积极方面,我们表明,如果预言机为其每个答案附加加密签名,那么假设仅存在抗碰撞哈希函数,每个预言机辅助计算都可以通过高效的证明者和验证者以零知识方式进行验证。除了隐私之外,这也提供了一种可扩展监督的替代方法,该方法既不依赖于辩论中的诚实对手,也不依赖于先前单证明者协议中计算的鲁棒性。
英文摘要
AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.