arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06051cs.CVcs.AI

图像尺度鲁棒性与视觉识别性能:跨架构分析

Image-Scale Robustness and Visual Recognition Performance: A Cross-Architecture Analysis

Anish Monsley Kirupakaran

AI总结:

本研究通过分析20个预训练ImageNet-1K分类器,发现尺度鲁棒性主要由基线识别精度决定,而非模型大小或架构家族,并提出了特征尺度作为量化指标。

AI中文摘要:

视觉识别模型对图像尺度变化的敏感性已得到充分证实,然而,在不同架构之间,控制这种敏感性的因素仍不清楚。在本工作中,我们研究了尺度鲁棒性是否在现代视觉模型之间表现出共同的定量结构。我们评估了20个预训练的ImageNet-1K分类器,涵盖七个架构家族,包括卷积、移动、高效和基于Transformer的架构。通过系统地降低输入图像尺度,我们构建了尺度-精度响应曲线,并将特征尺度定义为识别性能显著退化开始的紧凑度量。然后,我们考察了特征尺度与基线识别精度、模型参数数量、架构家族和表示稳定性之间的关系。观察到基线精度与特征尺度之间存在强负相关(Pearson r = -0.890,R^2= 0.792,p < 10^-6)。这种关系在自助重采样、留一架构分析和留一家族分析中保持稳定。相比之下,在控制基线精度后,参数数量提供的额外解释力可忽略不计(p = 0.80),而架构家族未提供显著的增量解释力。此外,特征尺度与表示稳定性之间基本没有关联(r = -0.003,p = 0.991)。这些结果表明,在所研究的模型中,尺度鲁棒性主要由基线识别性能组织,而非简单地由模型大小、架构家族或表示稳定性决定。本研究为表征跨视觉架构的尺度鲁棒性提供了一个实证框架,并识别出一种可复现的精度-尺度规律,值得进一步的理论研究。

英文摘要:

The sensitivity of visual recognition models to changes in image scale is well established, yet the factors governing this sensitivity across heterogeneous architectures remain unclear. In this work, we investigate whether scale robustness exhibits a common quantitative structure across modern vision models. We evaluate 20 pretrained ImageNet-1K classifiers spanning seven architectural families, including convolutional, mobile, efficient, and Transformer-based architectures. By systematically reducing input image scale, we construct scale-accuracy response curves and define a characteristic scale as a compact measure of the onset of substantial recognition degradation. We then examine the relationship between characteristic scale and baseline recognition accuracy, model parameter count, architectural family, and representation stability. A strong inverse association is observed between baseline accuracy and characteristic scale (Pearson r = -0.890, R^2= 0.792, p < 10^-6). This relationship remains stable under bootstrap resampling, leave-one-architecture-out analysis, and leave-one-family-out analysis. In contrast, parameter count provides negligible additional explanatory power after controlling for baseline accuracy (p = 0.80), while architectural family does not provide significant incremental explanatory power. Furthermore, characteristic scale shows essentially no association with representation stability (r = -0.003, p = 0.991). These results indicate that, across the studied models, scale robustness is strongly organized by baseline recognition performance rather than simply by model size, architectural family, or representation stability. The study provides an empirical framework for characterizing scale robustness across vision architectures and identifies a reproducible accuracy-scale regularity that warrants further theoretical investigation.

↑