AI 中文总结
本研究通过成对评估关联干净与扰动预测,证明类别级计数可界定可靠性界限,而图像特定合成响应能更有效地优先分配物理测试,提升失败召回率。
AI 中文摘要
鲁棒性评估必须考察多样的视觉扰动,而基准仅覆盖部分真实世界条件,物理测试成本高昂。成对评估将同一图像的干净预测与扰动预测关联起来,捕捉正确性、置信度和接受度在聚合准确率之外的变化。我们研究了这种图像对应关系如何支持鲁棒性评估中的两个需求:解释成对评估结果和为物理测试优先选择样本。为解释成对评估结果,我们固定两组预测记录,并在每个类别内改变它们的对应关系。我们证明,类别内的正确-正确计数在丢失接受度和保留正确输入的平均真实类别概率下降上,与任何可行的五状态细化给出相同的尖锐界限。区分持续错误与变化错误可进一步约束接受错误转换,而共享对应关系可建立独立成本区间未解决的政策排序。为优先选择物理测试样本,我们保留每个图像的合成响应,并按扰动下的平均真实类别概率对干净正确图像排序。在44个分类器中,测试最高风险20%的样本在轻度屏幕和打印重拍下分别发现67%和45%的失败,而干净置信度对应58%和36%,等规模自然变换平均对应60%和37%。结合概率平均和A3Rank评分适配,测试的扰动集比自然变换集产生更高的平均失败召回率;分数差异取决于来源和预算。这些发现共同表明,对应关系的价值取决于评估目标:类别级计数足以满足指定的可靠性界限,而图像特定的合成响应改善了评估池内物理测试的分配。
英文摘要
Robustness evaluation must examine diverse visual perturbations, while benchmarks cover only some real-world conditions and physical testing is costly. Paired evaluations link clean and perturbed predictions for the same image, capturing changes in correctness, confidence, and acceptance beyond aggregate accuracy. We investigate how this image correspondence supports two needs in robustness evaluation: interpreting paired evaluation results and prioritizing samples for physical testing. To interpret paired evaluation results, we fix both sets of prediction records and vary their correspondence within each class. We prove that classwise correct-correct counts give the same sharp bounds on lost acceptance and mean true-class probability decrease among retained-correct inputs as any feasible five-state refinement. Distinguishing persistent from changed wrong answers can further constrain accepted-error transitions, while shared correspondence can establish policy orderings left unresolved by separate cost intervals. To prioritize samples for physical testing, we retain each image's synthetic responses and rank clean-correct images by their mean true-class probability under corruption. Across 44 classifiers, testing the highest-risk 20% finds 67% and 45% of failures under mild screen and print recaptures, versus 58% and 36% for clean confidence and 60% and 37% for an equal-size natural-transformation average. With both probability averaging and an A3Rank scoring adaptation, the tested corruption set yields higher mean failure recall than the natural-transform set; differences between scores depend on the source and budget. Together, these findings show that the value of correspondence depends on the evaluation objective: classwise counts suffice for specified reliability bounds, while image-specific synthetic responses improve the allocation of physical tests within the evaluated pool.
Comments53 pages, 8 figures, including appendices