AI 中文总结
本文提出DISCERN协议,利用模型分歧无需标签的特性,通过零标签层和审计层实现认证风险差异审计,在保证有限样本有效性的同时,将标签复杂度降低1/rho因子,实验验证了高功效和低误报率。
AI 中文摘要
每个生产模型都会通过重新训练、微调、量化或供应商静默替换进行更新,而每次更新都可能比它所替代的模型更差。我们将更新晋升形式化为认证的配对风险差异审计。我们的出发点是支撑恒等式:两个模型之间的风险差异存在于它们不一致的输入上,无需标签即可观测。我们构建了DISCERN,一个顺序的两层协议。零标签层仅通过无标签流量认证分歧率低于容差阈值的良性更新。审计层通过任意时点有效的置信序列仅对采样的分歧进行标注,该置信序列在任何停止时间和任何标签路由规则下均有效,即使面对对抗性裁判也是如此。我们证明了有限样本有效性和匹配的标签复杂度界限,其阶为rho^2/eps^2(在速率水平上),因此利用免费分歧可证明比任何配对盲审计器节省1/rho的因子,并且该保证可跨无界序列的晋升从一个错误预算中组合。在785个更新对上的14,000多个重放审计流中,包括高达1.4B参数的语言模型的LoRA微调,误覆盖率为0.0002(名义5%),功效为0.986且零误报,56%的良性更新在零标签下获得认证。每次审计都会发出一个机器可检查的证据记录,用于上市后监测。
英文摘要
Every production model is updated, by retraining, fine-tuning, quantization, or a silent vendor swap, and each update risks being worse than what it replaced. We formalize update promotion as certified paired risk-difference auditing. Our starting point is a support identity: the risk difference between two models lives on the inputs where they disagree, observable without labels. We build DISCERN, a sequential two-tier protocol. A zero-label tier certifies benign updates whose disagreement rate is below tolerance from unlabeled traffic alone. An audited tier labels only sampled disagreements through an anytime-valid confidence sequence, valid at every stopping time and under any label-routing rule, even an adversarial judge. We prove finite-sample validity and matching label-complexity bounds of order rho^2/eps^2 at the rate level, so exploiting free disagreement provably saves a factor 1/rho over any pairing-blind auditor, and the guarantee composes across an unbounded sequence of promotions from one error budget. Across 14,000+ replayed audit streams over 785 update pairs, including LoRA fine-tunes of language models up to 1.4B parameters, miscoverage is 0.0002 (nominal 5%), power 0.986 with zero false alarms, and 56% of benign updates certify with zero labels. Each audit emits a machine-checkable evidence record for post-market monitoring.
Comments32 pages, 6 figures, 6 tables