AI 中文总结
针对评估器漂移下的模型发布认证,提出组合警觉方法,混合历史引导与当前标签投注,在保证有效性的同时显著减少标签需求,实验验证其高效性与稳健性。
AI 中文摘要
发布模型更新需要认证其当前总体风险低于阈值。可信标签成本高昂,而廉价评估器(如LLM裁判)可为每个示例打分。复用早期审计中的评估器错误颇具吸引力,但此类证据何时能替代当前标签?这取决于历史状态。若错误可能不可见地变化,则无标签测试无法检测该变化,且每个有效且有用的认证器必须持续以我们刻画的速度购买标签;若假设变化有界,则无标签认证在显式错误成本下有效。对于历史信息有用但不可信的中等情形,我们提出“组合警觉”(portfolio vigilance),一种顺序认证器,将受历史引导的投注专家与仅从当前标签学习的专家混合;历史仅影响其投注方式,因此有效性对任何历史均成立。贡献不在于先验信息投注或专家混合,而在于区分可进入有效性的历史与仅可指导标签收集的历史。在典型模型中,准确历史缩短决策时间但绝不提高证据增长率;过时历史可摧毁该增长率。在留出的CIFAR-10N和DICES-990数据上,组合警觉所需标签分别为匹配的预测增强监测器的0.465(95%置信区间[0.327,0.575])和0.740([0.618,0.877])倍,且未观察到错误认证,并在全部六个外部块上使用更少标签。在受损建议下,其保持在较优组件的8.0%以内,而仅信任历史的成本高达1.66倍。在DICES-990和ToxicChat上的事后确认重复裁判实验中,改变固定LLM裁判的评分标准使其分数超出运行间变异;此时组合所需标签分别为匹配监测器的0.790([0.667,0.909])和0.631([0.520,0.770])倍,且少于仅信任历史。
英文摘要
Releasing a model update requires certifying that its current-population risk stays below a threshold. Trusted labels are expensive, while a cheap evaluator, such as an LLM judge, scores every example. Reusing evaluator errors from earlier audits is tempting, but when may such evidence replace current labels? It depends on the status of history. If the errors can change invisibly, no label-free test detects the change, and every valid, useful certifier must keep buying labels at a rate we characterize; if a bound on the change is assumed, label-free certification is valid at an explicit error cost. For the middle ground, where history is informative but untrusted, we propose \emph{portfolio vigilance}, a sequential certifier mixing a betting expert guided by history with one that learns only from current labels; history affects only how it bets, so validity holds for any history. The contribution is not prior-informed betting or expert mixtures, but separating history that may enter validity from history that may only guide label collection. In a canonical model, accurate history shortens decisions but never raises the evidence growth rate; stale history can destroy it. On held-out CIFAR-10N and DICES-990 data, portfolio vigilance needs 0.465 (95\% CI $[0.327,0.575]$) and 0.740 ($[0.618,0.877]$) times the labels of a matched prediction-powered monitor, with no observed false certification, and fewer labels on all six external blocks. Under corrupted advice it stays within 8.0\% of its better component, while trusting history alone costs up to 1.66 times as much. In post-confirmatory repeated-judge experiments on DICES-990 and ToxicChat, changing a fixed LLM judge's rubric moves its scores beyond run-to-run variation; the portfolio then needs 0.790 ($[0.667,0.909]$) and 0.631 ($[0.520,0.770]$) times the labels of the matched monitor, and fewer than trusting history alone.
Comments27 pages, 8 figures