RiskBlend:机器学习回归测试中测试输入优先级排序的多信号框架
RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing
- East Carolina University(东卡罗来纳大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
RiskBlend是一种多信号优先级排序框架,融合四类跨版本风险信号,在1200种实验配置下于所有80种组合中均获最高平均APFD,可高效检测机器学习回归故障。
AI中文摘要:
当机器学习分类器被重新训练时,旧模型版本正确分类的输入可能会被更新版本错误分类,从而产生回归故障;由于验证预测结果是否与真实值一致可能需要人工标注、专家审查或昂贵的模拟(而非低成本的模型推理),这类故障的检测成本很高。测试输入优先级排序通过对输入进行排名,使得有限的验证预算能揭示尽可能多的回归故障,以此解决该问题。现有方法主要依赖单模型置信度分数,未利用模型版本间预测结果、决策边界及局部邻域的变化。我们提出RiskBlend,一种与分类器无关的优先级排序框架,它结合了四种互补的风险信号:历史故障模式、预测偏移、决策边界偏移和邻域变化;这些信号通过经验证学习的APFD-squared加权方法进行融合。在四个数据集、五个分类器、四个回归更新场景及15个随机种子(总计1200种实验配置)下,RiskBlend在全部80种数据集-分类器-场景组合中均达到最高平均APFD,相较于最强基准方法的APFD提升最高达0.32。基于置信度的方法仅在稀疏分类特征上的线性分类器中仍具竞争力,我们将此归因于特征空间几何特性。结果表明,跨版本行为信号为机器学习系统中的回归故障优先级排序提供了重要的互补信息。
英文摘要:
When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version, creating regression faults that are costly to detect because verifying predictions against ground truth may require human annotation, expert review, or expensive simulation rather than inexpensive model inference. Test input prioritization addresses this problem by ranking inputs so that a limited verification budget reveals as many regression faults as possible. Existing approaches rely predominantly on single-model confidence scores and do not exploit how predictions, decision boundaries, and local neighborhoods change between model versions. We propose RiskBlend, a classifier-agnostic prioritization framework that combines four complementary risk signals: historical failure patterns, prediction shift, decision-boundary shift, and neighborhood change. These signals are combined using validation-learned APFD-squared weighting. Across four datasets, five classifiers, four regression-update scenarios, and 15 random seeds, totaling 1,200 experimental configurations, RiskBlend achieves the highest average APFD in all 80 dataset-classifier-scenario combinations, with improvements of up to 0.32 APFD over the strongest baseline. Confidence-based methods remain competitive primarily for linear classifiers on sparse categorical features, which we attribute to feature-space geometry. The results show that cross-version behavioral signals provide important complementary information for prioritizing regression faults in machine learning systems.