针对欺骗性模型提供者的防操纵遗忘审计
Manipulation-Proof Oblivious Audits against Deceptive Model Providers
浏览论文内容
中文总结 AI 辅助
本文提出新型防操纵遗忘审计协议,利用私有信息检索机制提升审计对模型提供者操纵行为的可检测性,实验验证其有效实用。
中文摘要 AI 辅助
审计已成为算法治理的关键工具,为机器学习模型提供外部审查与治理机制,但确保此类评估的完整性仍是难题。例如在监管场景中,审计通常是预先声明的或易被检测,使模型提供者(无论有意或无意)可操纵审计过程,该漏洞在公平性评估中尤为突出——提供者常可推断敏感属性,策略性地均衡群体间分配率以满足公平性指标。本文提出一种新型审计协议,通过让审计方以遗忘方式查询模型,显著提升审计后对操纵行为的可检测性。该方法利用私有信息检索(Private Information Retrieval)机制,要求提供者对大量实例打标签,同时阻止其知晓最终将用于审计的子集。此协议高效,对审计方的开销极小,且无需修改被审计模型、其训练流程或推理管线。我们提供理论保证:在此协议下,试图隐藏不公平性的提供者必须伪造大量响应,从而提升操纵的难度与被检测概率。在代表性审计场景中的实验结果证实了该方法的有效性与实用性。
英文摘要
Audits have emerged as a critical instrument for algorithmic governance, providing a mechanism for external scrutiny and governance of machine learning models. However, ensuring the integrity of such assessments remains a challenging issue. For instance in regulatory contexts, audits are typically declared or easily detected, thus enabling model providers to manipulate the process, whether intentionally or inadvertently. This vulnerability is particularly acute in the context of fairness evaluations, in which providers can often infer sensitive attributes and strategically equalize allocation rates between groups to satisfy fairness metrics. In this paper, we introduce a novel audit protocol designed to significantly increase the post-audit detectability of such manipulations by enabling the auditor to query the model in an oblivious manner. Our approach leverages a Private Information Retrieval mechanism to require the provider to label a large set of instances, while preventing it from knowing which subset will ultimately be used for the audit. The protocol is efficient, imposes minimal overhead on the auditor, and requires no modification to the audited model, its training procedure, or its inference pipeline. We provide theoretical guarantees showing that, under this protocol, a provider attempting to hide unfairness must falsify a significantly larger number of responses, thereby increasing both the difficulty and the likelihood of detection of manipulation. Experimental results across representative audit scenarios confirm the effectiveness and practicality of our approach.