arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理论上可认证,实践中却被破解:密码学模型认证中的假设差距

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification

Carter Luck, Olive Franzese-McLaughlin, Elisaweta Masserova, Akira Takahashi, Antigoni Polychroniadou, Nicolas Papernot

arXiv 2607.21839首次发表:更新:

发表机构

University of Massachusetts Amherst; Vector Institute & University of Toronto; Carnegie Mellon University; J.P.Morgan AI Research & AlgoCRYPT CoE; University of Toronto(马萨诸塞大学阿默斯特分校; 向量研究所及多伦多大学; 卡内基梅隆大学; 摩根大通人工智能研究及AlgoCRYPT中心; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究密码学模型认证中现有安全定义与实际应用的差距,通过精心设计训练数据可攻击相关认证方案。为此形式化严格安全概念,引入通用协议模板,为设计安全隐私保护的机器学习审计协议提供指导。

AI 中文摘要

隐私保护机器学习审计协议允许审计师评估模型的准确性或公平性等属性,而不暴露其内部结构或训练数据,这使其对医疗或金融等敏感领域的模型审计很有吸引力。然而,现有安全定义往往达不到要求,大多数仅在固定审计数据集上认证模型行为,无法确保相同保证能推广到来自同一分布的其他数据集。我们表明,这种差距使模型提供者能通过精心设计训练数据攻击许多基于安全零知识证明的密码学模型认证(CMC)方案,导致模型在审计时表现良好,但在实践中行为异常。例如,攻击者能认证模型在审计数据集上准确率超99%,但在来自同一分布的新样本上准确率不到30%。为解决这一差距,我们为CMC框架形式化了严格的密码学安全概念,引入通用协议模板并证明其满足要求。我们的结果为现有方法提供了警示证据,并为设计安全的隐私保护机器学习审计协议提供了建设性指导。

英文摘要

Privacy-preserving machine learning auditing protocols allow auditors to assess models for properties such as accuracy or fairness, without revealing their internals or training data. This makes them especially attractive for auditing models deployed in sensitive domains such as healthcare or finance. For these protocols to be meaningful in real-world audit settings, though, their guarantees must reflect how the model will behave once deployed, rather than merely certifying its behavior during an audit. Existing security definitions often miss this mark: most certify model behavior only on a fixed audit dataset, without ensuring that the same guarantees generalize to other datasets drawn from the same distribution. As we show, this gap allows a model provider to attack many cryptographic model certification (CMC) schemes built on secure zero knowledge proofs (ZKP) by carefully engineering training data, resulting in models that exhibit benign behavior during an audit, but pathological behavior in practice. For example, we empirically demonstrate that an attacker can certify that a model achieves over 99% accuracy on an audit dataset, but less than 30% accuracy on fresh samples from the same distribution. To address this gap, we formalize rigorous cryptographic security notions tailored to CMC frameworks, introduce a generic protocol template, and prove that it satisfies these requirements. Our results thus offer both cautionary evidence about existing approaches and constructive guidance for designing secure, privacy-preserving ML auditing protocols.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑