码本引导的跨模态知识蒸馏用于结构异构特征
Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features
另 1 家 · 查看机构详情
- Kyungpook National University(庆北国立大学)
- AX/PI Center, Samsung Electronics(三星电子AX/PI中心)
- Chung-Ang University(中央大学)
- University of Seoul(首尔市立大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对结构异构跨模态特征难以对齐的问题,提出基于向量量化码本的跨模态蒸馏框架,以概念级锚点传递教师知识,在分类和语义分割任务上验证有效性。
中文摘要 AI 辅助
跨模态知识蒸馏将知识从教师模态转移到学生模态。现有的特征级对齐方法通常假设教师和学生特征位于结构上可对齐的表示空间中。然而,当跨模态特征在结构上异构且缺乏清晰的单元级对应关系时(例如二维空间视觉网格和一维时间音频序列),这一假设不成立,从而限制了特征级对齐的适用性。为解决这一挑战,我们提出了一种跨模态蒸馏框架,通过向量量化码本实现跨结构异构特征空间的有效知识转移。具体而言,教师特征被抽象为一组向量形式的码,无论其原始特征结构如何,所选码作为学生学习的概念级锚点。码的选择同时受任务相关性和学生兼容性的指导,使学生无需直接进行单元级特征对齐即可接收可转移的教师知识。在多种跨模态蒸馏场景中的实验结果表明,所提出的框架在分类和语义分割任务上具有有效性。
英文摘要
Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.