arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GRAFT:通过持续教师蒸馏成长聚合基础模型

GRAFT: Growing Agglomerative Foundation Models via Continual Teacher Distillation

Zhenghao Zhao, Chi Zhang, Qingshuang Chen, Yelin Kim

arXiv 2610.02597首次发表:更新:

发表机构

University of Illinois Chicago; Amazon(伊利诺伊大学芝加哥分校; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GRAFT提出持续多教师蒸馏框架,通过教师特定读出标记和几何无关关系损失,使单一骨干网络以单次蒸馏代价逐步统一图像理解、2D/3D预测及视觉-语言等多领域能力。

AI 中文摘要

视觉基础模型(如DINOv2、SigLIP2和MASt3R)从不同的预训练目标中发展出互补的能力,然而它们的知识仍分散在各自独立的专用模型中。多教师知识蒸馏为将这些能力整合到一个聚合骨干网络中提供了一条路径,但现有方法假设教师集合是固定的,而纳入新教师需要对整个教师集合重复进行昂贵的联合蒸馏。我们提出了GRAFT,一个持续的多教师蒸馏框架,使得统一的骨干网络能够从开放序列的基础模型中逐步获取能力。当新教师到来时,GRAFT将先前蒸馏得到的模型视为教师以保留已学能力,同时当前学生模型联合从先前模型和新教师中学习。此外,为了协调异构教师之间不兼容的表示几何,我们引入了教师特定读出标记,赋予每位教师对共享编码器的独立读出能力,以及几何无关关系损失,该损失通过匹配图像-文本相似性结构而非原始特征值来对齐视觉-语言教师。我们提供了GRAFT模型,这是一个单一的、可持续扩展的骨干网络,统一了五个领域,包括图像理解、2D稠密预测、3D人体姿态估计、3D视觉和视觉-语言,在所有领域均展现出强大性能,同时以单次蒸馏而非完全重新蒸馏的代价获取每个新能力。

英文摘要

Vision foundation models such as DINOv2, SigLIP2, and MASt3R develop complementary capabilities from different pretraining objectives, yet their knowledge remains distributed across separate, specialized models. Multi-teacher knowledge distillation offers a path toward consolidating these capabilities into a single agglomerative backbone, but existing approaches assume a fixed set of teachers, and incorporating a new teacher requires repeating expensive joint distillation over the entire teacher set. We introduce GRAFT, a continual multi-teacher distillation framework that enables a unified backbone to progressively acquire capabilities from an open-ended sequence of foundation models. When a new teacher arrives, GRAFT treats the previously distilled model as a teacher for preserving learned capabilities, while the current student jointly learns from both the previous model and the incoming teacher. Furthermore, to reconcile the incompatible representation geometries of heterogeneous teachers, we introduce Teacher Specific Readout Tokens, which grant each teacher an independent read-out of the shared encoder, together with Geometry Agnostic Relational Loss that aligns a vision-language teacher by matching image-text similarity structures rather than raw feature values. We provide GRAFT model, which is a single, continually extensible backbone that unifies five domains, including image understanding, 2D dense prediction, 3D human pose estimation, 3D vision, and vision-language, delivering strong performance across all of them while acquiring each new capability at the cost of a single distillation rather than a full re-distillation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑