arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13714cs.LG

以切片级非回归与在位者回退机制认证模型升级

Certifying Model Upgrades with Slice-Wise Non-Regression and Incumbent Fallback

Shengwei Zhang, Tao Wu, Fei Qian

首次发表
浏览论文内容

中文总结 AI 辅助

针对模型升级可能损害特定切片的问题,提出分离候选搜索与独立配对评估的发布流程,以非回归容差认证非劣效性,失败时回退在位者,并提供有限样本保证与实验验证。

中文摘要 AI 辅助

更新后的模型可能在改善总体指标的同时,使对下游用户至关重要的某个切片(数据子集)性能退化。我们研究在相对于保留的在位者(旧模型)的非回归容差约束下的检查点选择问题。核心区别在于“未能检测到损害”与“认证非劣效性”:前者在评估存在噪声时,可能以高概率发布有害的更新。我们提出一种可复现的发布流程,将候选搜索与独立的配对评估分离,并在认证失败时精确返回在位者。应用既有的交并原理与“先学习后测试”原则,我们为单个冻结候选、有限候选库及预定的测试顺序给出了有限样本保证。联合发布决策不需要按切片数量进行Bonferroni惩罚,尽管认证功效仍可能随切片数量增加而下降。在有界分数模拟中,一个“未检测到损害”的门控在某个32切片设置下,有99.7%的试验会发布有害候选,而精确的非劣效性门控在5%目标下仅为2.6%。一个构造的两块族在匹配候选数量下,比标量路径产生更大的认证效用。公开数字实验(包括后续提高平均总体准确率的延续实验)在每次运行中都返回在位者,因为认证功效不足。这些结果建立了一个可审计的协议及其局限性;它们并未证明对基础模型或多语言翻译升级的益处。

英文摘要

An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to non-regression tolerances relative to a retained incumbent. The central distinction is between failing to detect harm and certifying non-inferiority: the former can release harmful updates with high probability when evaluation is noisy. We give a reproducible release procedure that separates candidate search from independent, paired evaluation and returns the exact incumbent when certification fails. Applying established intersection-union and Learn-then-Test principles, we state finite-sample guarantees for one frozen candidate, a finite candidate library, and a prespecified testing order. A joint release decision does not require a slice-count Bonferroni penalty, although certification power can still decrease with the number of slices. In bounded-score simulations, a no-detected-harm gate releases a harmful candidate in 99.7% of trials in one 32-slice setting, compared with 2.6% for an exact non-inferiority gate at a 5% target. A constructed two-block family yields larger certified utility than a scalar path under matched candidate counts. Public digits experiments, including a subsequent continuation that improves average aggregate accuracy, return the incumbent in every run because certification is underpowered. These results establish an auditable protocol and its limitations; they do not establish benefits on foundation-model or multilingual translation upgrades.

发表机构

  • University of Pennsylvania(宾夕法尼亚大学)
  • Alibaba International Digital Commerce(阿里巴巴国际数字商业集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑