不破坏,而是验证:用于隐私保护大语言模型验证的对抗探针
Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification
浏览论文内容
中文总结 AI 辅助
该研究提出基于zk-SNARK的隐私保护审计框架,构建三类互补探针验证大语言模型,实验表明黑盒场景下的基于token的探针灵敏度最优,且Groth16 zk-SNARK工作流具备良好实用性。
中文摘要 AI 辅助
大语言模型部署后的变更可能改变其行为,同时常规输出却基本保持不变,当模型权重为专有信息时,这给AI治理带来了挑战。我们提出一种基于零知识简洁非交互式知识论证(zk-SNARK)的隐私保护审计框架,该框架旨在构建具有对抗样本特性的探针,以放大经核准模型与修改后部署版本之间的对数几率漂移。我们的框架在不同访问模型下探索互补的探针族:基于token的探针在黑盒场景下运行,仅需输入接口、分词器和词汇表;基于嵌入的探针需要对嵌入接口的灰盒访问权限;压力探针依赖额外的接口能力,但无需对模型权重或架构进行全白盒访问。这种多样性允许根据灵敏度、访问要求和部署成本来选择探针。我们在大语言模型架构、代表部署后攻击的模型篡改场景及GPU平台上评估探针构建。重要的是,实验结果表明,基于token的探针在黑盒场景下,在不同模型和GPU平台上始终提供最强的平均灵敏度。我们的Groth16 zk-SNARK工作流在探针集规模从1扩展到50时仍保持实用性:证明时间从1.02秒增加到1.78秒,验证时间保持在约0.84秒附近,证明大小保持恒定。
英文摘要
Post-deployment changes to large language models can alter behavior while leaving routine outputs largely unchanged, creating a challenge for AI governance when model weights are proprietary. We present a privacy-preserving zk-SNARK-based audit framework that searches for probes designed in the spirit of adversarial examples to amplify logit drift between an approved model and a modified deployment. Our framework explores complementary probe families under different access models. Token-based probes operate in a black-box setting and require only the input interface, tokenizer, and vocabulary. Embedding-based probes require gray-box access to the embedding interface. Stress probes rely on additional interface capabilities but do not require full white-box access to model weights or architecture. This range allows probe selection to balance sensitivity, access requirements, and deployment cost. We evaluate probe constructions across LLM architectures, model-tampering scenarios representative of post-deployment attacks, and GPU platforms. Importantly, our experimental results demonstrate that token-based probes consistently deliver the strongest mean sensitivity across models and GPU platforms, although operating in a black-box setting. Our Groth16 zk-SNARK workflow remains practical as the probe set scales from 1 to 50, where proving time increases from 1.02 to 1.78 seconds, verification remains near 0.84 seconds, and proof size remains constant.