Giga-Embeddings:用于高吞吐量文本嵌入的混合专家编码器
Giga-Embeddings: Mixture-of-Experts Encoders for High-Throughput Text Embeddings
- MIPT(莫斯科物理技术学院)
- SaluteDevices(萨卢特设备公司)
- MSU(莫斯科国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究推出Giga-Embeddings系列文本嵌入模型,含稀疏10B、密集3B、蒸馏480M三款模型,其10B模型在多语言等MTEB基准表现最优,吞吐量显著优于同类模型,且480M模型性能超FRIDA,参数更少。
AI中文摘要:
我们推出Giga-Embeddings系列文本嵌入模型,旨在将出色的检索质量与高效的推理服务相结合。该系列最大的成员是一个稀疏的100亿参数混合专家(Mixture-of-Experts,MoE)编码器,每个token约有18亿个激活参数。在英文、俄语、多语言及代码类的MTEB基准测试中,该模型在所有四个评估套件中均实现了系列内最强的综合性能。在输入为1024个token的vLLM基准测试中,它每秒处理114500个token,吞吐量比密集型30亿参数模型高25%,是所评估外部系统吞吐量的1.56至2.65倍。该系列还包含一个密集型30亿参数编码器和一个蒸馏后的4.8亿参数编码器,以适配更严格的计算和内存预算。我们使用与维度无关的目标函数训练该紧凑模型,该函数对齐教师模型与学生模型的相似度分布。最终的4.8亿参数模型在俄语MTEB上得分为70.98,超过了FRIDA,同时使用的参数减少了42%。我们发布了全部三个模型的检查点。
英文摘要:
We introduce Giga-Embeddings, a family of text embedding models designed to combine strong retrieval quality with efficient serving. Its largest member is a sparse 10B-parameter Mixture-of-Experts encoder with approximately 1.8B active parameters per token. Across English, Russian, multilingual, and code MTEB benchmarks, this model achieves the strongest aggregate performance within the family on all four evaluated suites. In our vLLM benchmark with 1024-token inputs, it processes 114.5k tokens per second, providing 25 percent higher throughput than the dense 3B model and 1.56-2.65x the throughput of the evaluated external systems. The family also includes a dense 3B encoder and a distilled 480M encoder for tighter compute and memory budgets. We train the compact model using a dimension-agnostic objective that aligns teacher and student similarity distributions. The resulting 480M model scores 70.98 on Russian MTEB, surpassing FRIDA while using 42 percent fewer parameters. We release all three model checkpoints.