发表机构
University of Zurich(苏黎世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究语言模型审计未发现泄漏时如何证明非泄漏,分析不同计算条件下的证明复杂度,并通过实验展示有限审计的遗漏率,区分计算条件与覆盖和执行条件。
AI 中文摘要
当语言模型审计未发现泄漏时,需要什么来证明非泄漏?我们研究在可执行的泄漏标准和解码规则下,对声明提示域所保证的性质。对于一般的多项式时间有界评估器,提供的泄漏执行是可多项式时间检查的,而泄漏存在性是\NP-完全的,确定性证明是\coNP-完全的。精确随机证明在每个固定有理截断点$(0,1)$处是$\coNP^{\PP}$-完全的。限制计算可以改变这些界限。例如,当所有随机性都是来自高效计算的有限概率表的终端抽取时,证明属于\coNP\\。注意力模型在局部依赖窗口为对数长度且先于一个全局头的情况下,在确定性解码、固定词汇表、精确有理加权均值、直接二元仿射读出和有限自动机提示域下,允许多项式时间证明。一个具有两个全局层的构造反而使证明在模板域上是\coNP-完全的,每层一个头,多项式宽度,对数精度和逆多项式logit边际。植入秘密实验衡量有限审计相对于完整参考所遗漏的内容。在30个秘密-模型状态对中,这些对在贪婪单提示执行下泄漏其秘密的4,096提示域,每对均匀选择256个记录评估会遗漏每个泄漏,期望为这些对的$41.06\\%$。批量检查和单提示检查在48个微调对中有一个完整域决策不一致,而相同顺序的重复再现了每个单提示输出。这些结果区分了证明的计算条件与解释阴性审计所需的覆盖和执行条件。
英文摘要
When a language-model audit finds no leak, what is needed to certify non-leakage? We study guarantees over a declared prompt domain under an executable leakage criterion and decoding rule. For general bounded polynomial-time evaluators, a supplied leaking execution is polynomial-time checkable, while leak existence is \NP-complete and deterministic certification is \coNP-complete. Exact stochastic certification is $\coNP^{\PP}$-complete at every fixed rational cutoff in $(0,1)$. Restricting the computation can change these bounds. For example, certification is in \coNP\ when all randomness is a terminal draw from an efficiently computed finite probability table. Attention models admit polynomial-time certification when local dependency windows of logarithmic length precede one global head, given deterministic decoding, fixed vocabulary, exact rational weighted means, a direct binary affine readout and finite-automaton prompt domains. A construction with two global layers instead makes certification \coNP-complete over template domains, with one head per layer, polynomial width, logarithmic precision and an inverse-polynomial logit margin. Planted-secret experiments measure what finite audits miss relative to complete references. Among 30 secret--model-state pairs that leak under greedy single-prompt execution on their secret's 4,096-prompt domain, uniformly selecting 256 recorded evaluations per pair misses every leak for an expected $41.06\%$ of these pairs. Batched and single-prompt checks disagree on one complete-domain decision among all 48 fine-tuned pairs, while a same-order repeat reproduces every single-prompt output. These results distinguish computational conditions for certification from the coverage and execution conditions needed to interpret a negative audit.