arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你可能在运行错误的 Inception 裁剪

You May Be Running the Wrong Inception Crop

Jason Chuan-Chih Chou

arXiv 2610.04147首次发表:更新:

发表机构

Cohere Labs Community(Cohere Labs 社区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现 TensorFlow/JAX 与 PyTorch 中 Inception 裁剪实现不同,通过大规模实验揭示裁剪尺度分布下尾决定增强强度,并验证 Beta 裁剪及结果普遍性。

AI 中文摘要

Inception 裁剪在提出十年后,已成为训练深度视觉模型的标准基于裁剪的数据增强方法。其均匀采样裁剪尺度和宽高比的做法不仅被广泛采用,其上下界也如此,其中尺度下界是唯一有时会被调整的例外。因此,令人惊讶的是,TensorFlow / JAX 生态系统中的标准实现以概率密度函数 $f(A) \propto \frac{1}{\sqrt{A}}$ 采样裁剪尺度,而 PyTorch 对应实现则遵循原始描述。受此发现启发,我们在 ImageNet-1k 数据集上使用各种训练预算和裁剪尺度分布训练了 522 个 ViT-S/16 模型。在 90 个 epoch 的训练预算下,我们达到了 $78.78\pm0.09$ 的 top-1 验证准确率,并发现:1. 更高的训练预算需要更强的增强;2. 裁剪尺度分布的下尾决定了 Inception 裁剪的增强强度;3. 使用更高训练预算训练的模型表现出更稀疏的显著性,无论裁剪尺度分布或权重衰减如何。基于发现 2,我们重新审视了 Beta 裁剪的性能,其较柔和的截止允许其在训练预算间以更小的妥协优化模型性能。我们使用 Scion 优化器和 AdamW 复制了发现 1 和 3,表明结果可能具有普遍性。

英文摘要

A decade after its inception, Inception crop has become the standard crop-based data augmentation method for training deep vision models. Not only is its practice of uniformly sampling crop scale and aspect ratio widely adopted, but also its lower and upper bounds, with the scale lower bound being the sole exception that is sometimes tuned. It is therefore surprising that the standard implementation in the TensorFlow / JAX ecosystem samples crop scale with probability density function $f(A) \propto \frac{1}{\sqrt{A}}$ unlike the PyTorch counterpart, which follows the original description. Motivated by this discovery, we train 522 ViT-S/16 models on the ImageNet-1k dataset with various training budgets and crop scale distributions. We reach $78.78\pm0.09$ top-1 val. accuracy with 90 epochs of training budget and find that 1. Higher training budget requires stronger augmentation; 2. Lower tail of the distribution of the crop scale determines the augmentation strength of Inception crop; 3. Models trained with higher training budget exhibit sparser saliency, regardless of the crop scale distribution or weight decay. Based on 2. we revisit the performance of Beta crop, whose softer cutoff allows it to optimize model performance across training budgets with less compromise. We replicate 1. and 3. with Scion optimizer in addition to AdamW, suggesting that the results may be general.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑