面向边缘视觉语言模型的高效量化感知蒸馏与跨模态对齐
Efficient Quantization-Aware Distillation with Cross-Modal Alignment for Edge Vision-Language Models
浏览论文内容
中文总结 AI 辅助
针对边缘设备上视觉语言模型部署,提出统一教师锚定框架联合优化蒸馏与量化,并设计交叉注意力适配器增强非RGB模态,实验证明有效。
中文摘要 AI 辅助
大规模视觉语言模型(VLM)如CLIP能够实现强大的开放词汇推理,然而在资源受限的边缘设备上部署这些能力仍然具有挑战性。EdgeVL通过将CLIP表示蒸馏到轻量级多模态编码器,并应用量化感知训练(QAT)在边缘硬件上实现高效开放词汇分类(OVC)来解决这一问题。然而,其两阶段优化对蒸馏和QAT采用了不同的目标,且对比学习在量化学生空间内进行,这可能导致优化不一致和训练效率降低。此外,RGB和非RGB模态采用相同的监督可能导致模态不平衡。我们提出了一种面向边缘部署的量化语义蒸馏统一框架。通过在统一的教师锚定框架内联合优化蒸馏和量化,我们的方法确保了量化下训练的一致性,抑制难负样本并扩大决策边界。此外,我们设计了一个轻量级交叉注意力适配器,通过RGB引导的语义迁移增强非RGB表示,缩小模态差距。大量实验表明,在保持部署效率的同时,非RGB模态上取得了持续改进。
英文摘要
Large-scale vision-language models (VLM) such as CLIP enable strong open-vocabulary reasoning, yet deploying these capabilities on resource-constrained edge devices remains challenging. EdgeVL addresses this problem by distilling CLIP representations into lightweight multi-modal encoders and applying quantization-aware training (QAT) for efficient Open-Vocabulary Classification (OVC) on edge hardware. However, its two-stage optimization applies different objectives for distillation and QAT, and contrastive learning is performed within the quantized student space, which can result in inconsistent optimization and reduced training efficiency. Moreover, identical supervision across RGB and non-RGB modalities may lead to modality imbalance. We propose a unified framework for quantized semantic distillation tailored to edge deployment. By jointly optimizing distillation and quantization within a unified teacher-anchored framework, our method ensures consistent training under quantization, suppressing hard negatives and enlarging decision margins. Additionally, we design a lightweight cross-attention adapter that enhances non-RGB representations through RGB-guided semantic transfer, narrowing the modality gap. Extensive experiments demonstrate consistent improvements on non-RGB modalities while maintaining deployment efficiency.
发表机构
- Korea University(高丽大学)
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。