arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32007cs.CC

模型评估的邻近性交互式证明

Interactive Proofs of Proximity for Model Evaluation

Geoffroy Couteau, Nikolas Melissaris, Tamara Paris

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出模型评估的邻近性交互式证明,实现双亚线性资源下认证统计属性,改进查询复杂度并应用于多种审计场景。

中文摘要 AI 辅助

我们研究用于模型评估的邻近性交互式证明(IPPs),其中资源受限的验证者与不可信的证明者(通常是模型所有者)交互,以在未知输入分布下认证模型的统计属性。我们的公式将输入分布的采样与查询模型及评估其输出分开;区分真实审计数据(黑盒采样)与生成数据(对采样器的选择随机性或灰盒访问);并允许证明者和验证者使用不同的评估器。我们专注于双亚线性IPPs,其中验证者和诚实证明者都使用亚线性资源,以及(加权)汉明权重属性。对于普通汉明权重,我们给出一个容错的双亚线性IPP。对于完备性和健全性半径$\varepsilon_c<\varepsilon_f$以及间隙$g=\varepsilon_f-\varepsilon_c$,一个对数轮次的实例化使用$\widetilde{O}(1/g)$次验证者查询和$O(1/g^2)$次诚实证明者查询,改进了Amir、Goldreich和Rothblum(ITCS 2025)的三次依赖。我们证明了匹配的查询下界,直到多对数因子。对于分布加权的汉明权重,黑盒采样需要$\Theta(1/g^2)$个验证者样本,但仅需$\widetilde{O}(1/g)$次评估;二次样本复杂度在内部区域是必要的。通过选择随机性访问,问题简化为普通汉明权重,产生$\widetilde{O}(1/g)$次调用和评估。如果双方的评估器在分布的$\rho$部分任意不一致,并在其他地方最多相差$\gamma$,则只要$g>2\kappa$,其中$\kappa=\rho+(1-\rho)\gamma$,我们的协议保持双亚线性。应用包括审计准确性、群体公平性、校准、无害性、有用性和平均情况鲁棒性。

英文摘要

We study interactive proofs of proximity (IPPs) for model evaluation, where a resource-limited verifier interacts with an untrusted prover, typically the model owner, to certify statistical properties of a model under an unknown input distribution. Our formulation separates sampling the input distribution from querying the model and evaluating its output; distinguishes real audit data (black-box sampling) from generated data (chosen-randomness, or gray-box, access to the sampler); and allows the prover and verifier to use different evaluators. We focus on doubly-sublinear IPPs, where both the verifier and honest prover use sublinear resources, and on (weighted) Hamming weight properties. For ordinary Hamming weight, we give a tolerant doubly-sublinear IPP. For completeness and soundness radii $\varepsilon_c<\varepsilon_f$ and gap $g=\varepsilon_f-\varepsilon_c$, a logarithmic-round instantiation uses $\widetilde{O}(1/g)$ verifier queries and $O(1/g^2)$ honest-prover queries, improving the cubic dependence of Amir, Goldreich, and Rothblum (ITCS 2025). We prove matching query lower bounds up to polylogarithmic factors. For distribution-weighted Hamming weight, black-box sampling requires $Θ(1/g^2)$ verifier samples but only $\widetilde{O}(1/g)$ evaluations; the quadratic sample complexity is necessary in the interior regime. With chosen-randomness access, the problem reduces to ordinary Hamming weight, yielding $\widetilde{O}(1/g)$ calls and evaluations. If the parties' evaluators disagree arbitrarily on a $ρ$-fraction of the distribution and by at most $γ$ elsewhere, our protocols remain doubly sublinear whenever $g>2κ$, where $κ=ρ+(1-ρ)γ$. Applications include auditing accuracy, group fairness, calibration, harmlessness, usefulness, and average-case robustness.

发表机构

  • Université Paris Cité(巴黎西岱大学)
  • CNRS(法国国家科学研究中心)
  • IRIF(IRIF研究所)
  • McGill University(麦吉尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑