“我能在哪里信任你?”:代理模型保真度的边界感知评估
"Where Can I Trust You?": Boundary-Aware Evaluation of Surrogate Fidelity
浏览论文内容
中文总结 AI 辅助
本研究探究代理模型与教师模型的不一致性是否集中于教师决策边界附近,发现边界附近保真度显著更低,且基于置信度区域的条件化一致性比全局分数更能预测评估集变化时的表现,但边界稳定性受深度教师训练随机性影响。
中文摘要 AI 辅助
代理模型通常通过它们在评估集上与教师模型的一致性频率来评估。这种一致性的局部变化是众所周知的,但其结构和后果尚不清楚。我们探究不一致性是否系统性地集中在教师模型的决策边界附近,以及保留这种结构是否提供了超越单一全局分数的信息。在多个数据集和代理模型类别上,我们发现在两种不同的识别近边界样本的方法下,教师决策边界附近的保真度显著较低。此外,基于置信度定义的区域对一致性进行条件化,相对于仅使用全局分数,能更好地预测当评估集组成变化时教师与代理模型的一致性。然而,那些在全局和决策边界附近都与教师模型同等一致的代理模型,对使用代理模型自身选择的更改可能反应截然不同。最后,我们表明,独立训练的深度教师模型可能在大多数预测上一致,但在识别哪些样本位于其决策边界附近时却可能不同,这使得将这些边界用作评估代理模型的稳定参考区域变得复杂。综合来看,这些结果表明,代理模型的保真度不仅取决于代理模型与教师模型的一致性频率,还取决于这种一致性在何处成立,并且对于深度模型而言,还取决于教师模型的决策边界在多次训练运行中的稳定性。
英文摘要
Surrogate models are commonly evaluated by how often they agree with their teacher model over an evaluation set. Local variation in this agreement is well known, but its structure and consequences are less clear. We ask whether disagreement is systematically concentrated near the teacher's decision boundary and whether retaining that structure provides information beyond a single global score. Across several datasets and surrogate model classes, we find substantially lower fidelity near teacher decision boundaries under two different methods of identifying near-boundary examples. Moreover, conditioning agreement on confidence-defined regions improves prediction of teacher--surrogate agreement when evaluation-set composition changes, relative to the global score alone. Yet surrogates that agree equally well with the teacher both globally and near the decision boundary can respond very differently to changes selected using the surrogate itself. Finally, we show that independently trained deep teachers can agree on most predictions while identifying different examples as lying near their decision boundaries, complicating the use of those boundaries as stable reference regions for evaluating surrogates. Together, these results show that surrogate fidelity depends not only on how often a surrogate agrees with its teacher, but also on where that agreement holds and, for deep models, how stable the teacher's decision boundary is across training runs.