Eddy-VL 1.9B:用于边缘可部署多模态嵌入的结构剪枝与分层蒸馏
Eddy-VL 1.9B: Structural Pruning and Layered Distillation for Edge-Deployable Multimodal Embedding
浏览论文内容
中文总结 AI 辅助
介绍用于边缘可部署视觉语言检索的Eddy-VL 1.9B模型,采用探针驱动结构剪枝和分层知识蒸馏压缩,参数减少约9.5%,在MMEB-V2等评估中保留教师模型大部分性能,降低延迟,证明其在受限边缘部署下多模态检索的有效性。
中文摘要 AI 辅助
在本报告中,我们介绍了Eddy-VL 1.9B,这是一个基于Qwen3-VL-Embedding-2B构建的压缩多模态嵌入模型,用于离线、边缘可部署的视觉语言检索。Eddy-VL针对云API不可用且低延迟至关重要的气隙取证和调查场景。压缩方法包括:(i)探针驱动的结构剪枝,通过相邻层线性CKA去除四个冗余文本解码器层(从28层减至24层);(ii)分层知识蒸馏,采用孔覆盖师生映射、中间层注意力图1-CKA以及最终层的MSE和余弦损失,维度为{128, 256, 512, 1024, 2048}。发布的模型包含1,926,188,032个参数(3.85 GB bf16),比21.3亿参数的教师模型少约9.5%。在MMEB-V2(78个任务,VLM2Vec协议)上的实证评估表明,Eddy-VL的总体得分为63.2,教师模型为68.9,在保留教师模型91.7%性能的同时,弥补了因剪枝单独损失的12.1分中的6.4分(56.8)。在SugarCrepe、MR2-Bench和ARO上的组合推理性能与教师模型相近,而Winoground组性能仍是主要限制。深度剪枝还将前向延迟降低了约10%。我们展示了架构、压缩方法、训练过程和评估结果,证明了Eddy-VL在受限边缘部署下进行多模态检索的有效性。模型权重和推理代码可在Hugging Face上公开获取。
英文摘要
In this report, we introduce Eddy-VL 1.9B, a compressed multimodal embedding model built on Qwen3-VL-Embedding-2B for offline, edge-deployable vision-language retrieval. Eddy-VL targets air-gapped forensic and investigative settings where cloud APIs are unavailable and low latency is essential. Compression combines (i) probe-driven structural pruning that removes four redundant text-decoder layers (28 to 24) ranked by adjacent-layer linear CKA, and (ii) layered knowledge distillation with hole-covering teacher-student mappings, mid-layer attention-map 1-CKA, and final-layer MSE and cosine losses with Matryoshka dimensions {128, 256, 512, 1024, 2048}. The released model contains 1,926,188,032 parameters (3.85 GB bf16), representing approximately 9.5% fewer parameters than the 2.13B teacher model. Empirical evaluations on MMEB-V2 (78 tasks, VLM2Vec protocol) show that Eddy-VL achieves an overall score of 63.2 compared with 68.9 for the teacher, retaining 91.7% of the teacher's performance while recovering 6.4 of the 12.1 points lost through pruning alone (56.8). Compositional reasoning performance remains close to the teacher on SugarCrepe (86.1 vs. 86.4), MR2-Bench (24.5 vs. 24.7), and ARO (59.5 vs. 60.4), while Winoground group performance (6.8 vs. 8.5) remains the primary limitation. Depth pruning also reduces forward latency by approximately 10% (150.0 to 136.4 ms per image on NVIDIA DGX Spark using FlashAttention-2). We present the architecture, compression methodology, training procedures, and evaluation results, demonstrating the effectiveness of Eddy-VL for multimodal retrieval under constrained edge deployment. Model weights and inference code are publicly available on Hugging Face.
发表机构
- Urock-AI Lab(Urock人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。