重新思考工业密集预测中的数据效率:决定视觉Transformer(ViTs)低数据优势的是预训练一致性,而非归纳偏置
Rethinking Data Efficiency in Industrial Dense Prediction: Pretraining Coherence, Not Inductive Bias, Determines ViTs Low-Data Advantage
- ZTE Corporation(中兴通讯股份有限公司)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究通过实验发现工业密集预测中ViTs的低数据优势源于预训练一致性而非归纳偏置,提出AlignBlock解决跨架构特征差距,明确数据效率边界并验证了嫁接颈部网络的性能提升。
AI中文摘要:
人们普遍认为,在工业密集预测任务中,视觉Transformer(ViTs)比卷积神经网络(CNNs)需要更多的标注数据。通过在四个工业数据集上开展受控实验,研究人员发现数据效率差距源于预训练不一致性,即ImageNet预训练的ViT骨干网络与COCO预训练的CNN颈部网络之间存在统计不匹配,而非ViTs固有的自注意力缺陷。研究人员表征了跨架构特征差距,并提出了轻量级AlignBlock系列用于金字塔级特征重校准。核心发现通过实验确定了数据效率边界:对于样本量≥200的领域邻近场景,Swin-Graft超越YOLOv11x(703样本时最终mAP@50为0.973,对比YOLOv11x的0.956);对于领域遥远场景,CNNs仍保持优势(141样本时钩子模型mAP@50为0.900,对比0.600)。经嫁接的颈部网络权重的mAP可达随机初始化颈部网络的2.5倍。
英文摘要:
Vision Transformers (ViTs) are widely believed to require more labeled data than CNNs for industrial dense prediction. Through controlled experiments on four industrial datasets, we show that the data-efficiency gap stems from pretraining incoherence, which refers to the statistical mismatch between ImageNet-pretrained ViT backbones and COCO-pretrained CNN necks, rather than from inherent self-attention deficits. We characterize the cross-architecture feature gap and propose a lightweight AlignBlock family for pyramid-level feature recalibration. Our core finding empirically identifies a data-efficiency frontier: for domain-proximal scenes with >= 200 samples, Swin-Graft surpasses YOLOv11x (terminal 703-shot: 0.973 vs 0.956 mAP@50); for domain-distant scenes, CNNs retain advantage (hook 141-shot: 0.900 vs 0.600 mAP@50). Grafted neck weights yield up to 2.5x the mAP of a randomly initialized neck.