arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估漂移下的模型重训练:累积子组差异的配对比较

Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity

Aaron Ceross

arXiv 2609.09788首次发表:更新:

发表机构

University of Birmingham(伯明翰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过配对比较评估了不同重训练策略在数据漂移下对累积子组差异的影响,发现所有策略均能降低差异,但比较结果依赖于评估总体和测量方式。

AI 中文摘要

选择何时对已部署的分类器进行重训练,需要评估所用模型序列中的子组错误率,包括更新之间的时段。我们将完整的计划策略、损失触发策略和子组差距触发策略与在相同观测和延迟标签上保留初始模型进行了比较。分别针对真阳性率和假阳性率,结果是部署窗口内绝对子组差距的配对差异之和。模拟中的总体评估、行动记录和替代调度评估了测量和重训练行为如何影响这些比较。在两种模拟漂移机制下,每个条件有400条新轨迹的后续样本中,所有三种策略的平均累积差异均较低,相当于每个窗口的平均差距减少了0.04到0.88个百分点。将未改变的模型与已知的生成分布进行评估,保留了所有平均方向,但有限窗口和总体比较在69%到92%的轨迹中一致认为更新是增加、减少还是保持累积差异不变。在子组特定漂移下,较小的真阳性率差距伴随着两个组中较低的敏感性。在探索性的美国社区调查重放中,人员加权逆转了所有三个种族假阳性率平均比较,而不改变预测或行动;所有三个加权区间都包含零。策略比较需要特定组别的比率、行动分布和明确的评估总体以及平均差异。这些分析是非确认性的。共享重放需要策略无关的观测和指定延迟后的完整标签。

英文摘要

Choosing when to retrain a deployed classifier requires assessing subgroup error rates across the sequence of models used, including periods between updates. We compare complete scheduled, loss-triggered, and subgroup-gap-triggered policies with retaining the initial model on the same observations and delayed labels. For true-positive and false-positive rates separately, the outcome is the paired difference in absolute subgroup gaps summed over deployment windows. Population evaluation in simulation, action records, and alternative schedules assess how measurement and retraining behaviour affect these comparisons. In a follow-up sample of 400 new trajectories per condition across two simulated drift regimes, all three policies had lower mean cumulative disparity, equivalent to reductions of 0.04 to 0.88 percentage points in the average gap per window. Evaluating the unchanged models against the known generating distributions preserved all mean directions, but finite-window and population comparisons agreed on whether updating increased, reduced or left cumulative disparity unchanged in 69 to 92 percent of trajectories. Under subgroup-specific drift, smaller true-positive-rate gaps accompanied lower sensitivity in both groups. In an exploratory American Community Survey replay, person weighting reversed all three race false-positive-rate mean comparisons without changing predictions or actions; all three weighted intervals included zero. Policy comparisons require group-specific rates, action distributions, and an explicit evaluation population alongside mean disparity. These analyses are non-confirmatory. Shared replay requires policy-independent observations and complete labels after the specified delay.

Comments44 pages, 10 figures, 34 tables. Reproducibility package version 2026.09.08-r2

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑