发表机构
Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过分析视觉Transformer的注意力结构,发现注意力迁移的差距源于特征而非注意力结构,且鲁棒性成熟晚于准确率,早停规则会对鲁棒性采样不足。
AI 中文摘要
近期研究表明,经训练以复制预训练教师模型注意力图的视觉Transformer(ViTs),在分布内任务中能恢复微调模型的大部分准确率,但在分布偏移场景下仍存在明显差距。此前从未直接测量过这种迁移在注意力结构中的表现,并将其与鲁棒性关联起来。我们针对ImageNet-100上自监督教师模型的ViT-S学生模型构建了相关分析工具,得出三项发现以支撑一个结论:其一,这种迁移本质上是完全且持久的:蒸馏得到的学生模型注意力与教师模型的相似度,约比微调模型高两个数量级,且不会随额外训练发生漂移;其二,在参数减少14倍、数据量减少10倍的场景下,该差距确实存在,但它与训练成熟度相关:差距随训练成熟度变化,若完成早停规则中断的训练计划,在三个随机种子中有两个可将差距缩小至我们预先注册的阈值以下,且在准确率相等的对比中也得出相同结果。该规模下的最终差距主要是训练成熟度的人为产物:鲁棒性的成熟晚于准确率,针对准确率调整的早停规则会对鲁棒性采样不足;其三,在两种经注册的准确率匹配方式下,将蒸馏与微调条件间的结构分离度减半以降低跨行冗余,未产生可检测的鲁棒性变化。经验证的注意力迁移、结构未发生变化但差距缩小,以及直接干预下的无效结果,共同表明该差距存在于特征层面,而非可见的注意力结构中。这是结合排除法与干预法的结论,其适用范围为我们测量的场景:在该场景下,注意力叠加图仅显示模型的注视位置,而非模型的知识。
英文摘要
Vision transformers (ViTs) trained to copy a pretrained teacher's attention maps recover most of fine-tuning's in-distribution accuracy yet fall measurably short of it under distribution shift, as recent work has shown. What the copy delivers has never been measured directly in the attention structure and tied to robustness. We build that instrumentation for ViT-S students of a self-supervised teacher on ImageNet-100, and report three findings that triangulate one conclusion. First, the transfer is essentially perfect and permanently so: the distilled student's attention ends up roughly two orders of magnitude closer to the teacher's than fine-tuning does, and does not drift with additional training. Second, the gap is real at 14$\times$ fewer parameters and 10$\times$ less data than previously studied, but it has a time axis. It tracks training maturity, and completing the schedules that the stopping rule interrupted closes it below our pre-registered threshold in two of three seeds, with comparisons at equal accuracy giving the same result. The endpoint gap at this scale is substantially a training-maturity artifact: robustness matures later than accuracy, and stopping rules tuned to accuracy undersample it. Third, forcing cross-row redundancy down by half the structural separation between the distilled and fine-tuned conditions produces no detectable robustness response under two registered ways of matching accuracy. Verified transfer, a gap that closes while the structure never moves, and a null under direct intervention are together consistent with the deficit residing in features, not in the visible attention structure. This is elimination plus intervention, and its scope is the regime we measured. In this regime, attention overlays show where a model looks, not what it knows.
Comments28 pages, 9 figures