ViTok:在AM-RADIO风格多教师蒸馏中利用PHI-S和掩码图像建模提升密集语义
ViTok: Improving Dense Semantics in AM-RADIO-Style Multi-Teacher Distillation with PHI-S and Masked Image Modelling
浏览论文内容
中文总结 AI 辅助
本文提出ViTok,通过整合适配器拆分、非对称损失、MIM及PHI-S平衡等修改,在多教师蒸馏中兼顾全局与密集语义,在ImageNet-1K和ADE20K上均达到或超越教师性能。
中文摘要 AI 辅助
我们研究如何将当前的VITOK进展整合为一个统一的多教师蒸馏方案,该方案同时保留全局识别和密集语义。我们的起点是一个从SigLIP2和DINOv3-L蒸馏得到的AM-RADIO风格学生模型,其中SigLIP2提供强大的全局语义,DINOv3-L提供更强的密集特征。核心实证问题是同一方案无法同等优化所有目标:改善ImageNet-1K kNN准确率的改动仍可能降低ADE20K分割性能。我们总结了一系列使这一权衡更明确且更易管理的修改:为CLS和patch token拆分适配器头、非对称余弦/MSE损失、从DINOv3-L检查点初始化、教师重新加权、掩码图像建模(MIM)以及PHI-S特征平衡。所得模型在ImageNet-1K kNN分类上达到83.2 patch kNN和85.2 CLS kNN,略超DINOv3-L教师,而PHI-S将ADE20K性能从46.5/58.1恢复至48.5/61.0 mIoU/mAcc,在该密集基准上与教师匹配。我们还总结了负面结果:将蒸馏从ImageNet-1K扩展到ImageNet22K并非始终有益,且天真地添加额外教师(如SAM3或HOG特征)会引入干扰。本文并非宣称最终方案,而是将当前项目状态提炼为紧凑的实证故事和一套具体教训,供未来迭代参考。
英文摘要
We study how to consolidate the current VITOK progress into a single multi-teacher distillation recipe that jointly preserves global recognition and dense semantics. Our starting point is an AM-RADIO-style student distilled from SigLIP2 and DINOv3-L, where SigLIP2 supplies strong global semantics and DINOv3-L supplies stronger dense features. The central empirical issue is that the same recipe does not optimize all objectives equally well: changes that improve ImageNet-1K kNN accuracy can still degrade ADE20K segmentation. We summarize a progression of modifications that make this trade-off more explicit and more manageable: split adaptor heads for CLS and patch tokens, asymmetric cosine/MSE losses, initialization from a DINOv3-L checkpoint, teacher reweighting, masked image modeling (MIM), and PHI-S feature balancing. The resulting model reaches 83.2 patch kNN and 85.2 CLS kNN, slightly surpassing the DINOv3-L teacher on ImageNet-1K kNN classification, while PHI-S restores ADE20K performance from 46.5/58.1 to 48.5/61.0 mIoU/mAcc, matching the teacher on this dense benchmark. We also summarize negative results: scaling distillation from ImageNet-1K to ImageNet22K does not consistently help, and naively adding extra teachers such as SAM3 or HOG features introduces interference. Rather than claiming a final recipe, this paper distills the current project state into a compact empirical story and a concrete set of lessons for future iterations.
发表机构
- Beihang University(北京航空航天大学)
- Indian Institute of Technology(印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。