arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26081cs.LGcs.CV

间隔下降坐标用于跨预算鲁棒性评估

Margin-Drop Coordinates for Cross-Budget Robustness Evaluation

Yanliang Huang, Zhen Zhang, Peng Xie, Wenyuan Wu, Sitong Zhu, Zhuoqi Zeng, Amr Alanwar

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出间隔下降坐标分解方法,利用浅层评估信息预测跨预算鲁棒性崩溃,在42个编码器上验证了其排序能力,优于传统存活率指标。

中文摘要 AI 辅助

固定预算的鲁棒性评估可能会选择错误的冻结视觉编码器。一个在浅层攻击下幸存的编码器,在相同评估被加强时,可能会失去大部分鲁棒性。我们探究浅层评估是否包含足够的信息来识别这种预算脆弱性。对于每个干净且正确分类的样本,评估记录了干净成对间隔、一阶线性化间隔下降尺度、从干净起点一步攻击的间隔下降,以及迭代攻击达到的间隔下降。通过该尺度归一化,得到三个间隔下降坐标,分别捕捉干净间隔余量、一步不足和漂移,其中漂移是迭代攻击超过一步扰动所达到的额外归一化间隔下降。三者共同重构了归一化攻击后间隔,从而决定通过或失败的结果。在42个预训练冻结视觉编码器中,浅层存活率对后续PGD-10到PGD-200的崩溃几乎不携带排序信息,Spearman相关系数为-0.006,而中位浅层漂移坐标对相同崩溃的排序为+0.811。该结果在保留的编码器池和ℓ∞评估下依然成立。在深度评估仅限于11个编码器时,根据浅层漂移排序恢复了17个高崩溃编码器中的11个,而根据存活率排序仅恢复5个。完整的坐标分解进一步区分了具有相同固定预算残差但在更深预算下分化的案例,并在干预下将间隔修复与漂移修复分离,揭示了端点鲁棒性单独无法识别的不同修复路径。

英文摘要

Fixed-budget robustness evaluation can select the wrong frozen vision encoder. An encoder that survives a shallow attack may lose most of that robustness when the same evaluation is strengthened. We ask whether the shallow evaluation contains enough information to identify this budget fragility. For each clean-correct sample, the evaluation records the clean pairwise margin, the first-order linearized margin-drop scale, the margin drop from a clean-start one-step attack, and the drop reached by an iterative attack. Normalizing by that scale gives three margin-drop coordinates capturing clean margin slack, one-step shortfall, and drift, where drift is the additional normalized margin drop the iterative attack reaches beyond the one-step perturbation. Together, they reconstruct the normalized post-attack margin and therefore the pass-or-fail outcome. Across 42 pretrained frozen vision encoders, the shallow survival rate carries essentially no rank information about subsequent PGD-10 to PGD-200 collapse, at Spearman -0.006, while the median shallow drift coordinate ranks the same collapse at +0.811. The result persists in a held-out encoder pool and under an $\ell_\infty$ evaluation. With deep evaluation limited to 11 encoders, ranking by shallow drift recovers 11 of the 17 high-collapse encoders, compared with 5 under survival-rate ranking. The full coordinate decomposition further distinguishes cases that share the same fixed-budget residual but diverge at deeper budgets, and separates margin repair from drift repair under interventions, revealing distinct repair paths that endpoint robustness alone does not identify.

↑