arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Bekko嵌入:使用超紧凑型编码器的参数高效多语言检索

Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders

Yuichi Tateno

arXiv 2607.25180首次发表:更新:

AI 中文总结

研究如何让多语言检索模型更小,提出Bekko Embedding,通过剪枝mmBERT - small并分两阶段训练,最小模型a8m参数少速度快,在多语言检索任务中表现出色,还发布相关资源支持可重复研究。

AI 中文摘要

具有竞争力的多语言检索模型能有多小?我们提出了Bekko嵌入:其最小的模型bekko - embedding - v1 - a8m只有不到800万个活动参数(AP)——主导推理计算的非嵌入参数,然而在官方MMTEB多语言v2检索(nDCG@10)中得分56.2,高于我们比较中的多语言 - e5系列和BGE - M3(AP的40倍)。质量更高的bekko - embedding - v1 - a25m(略低于2500万个AP)达到57.5,与gte - 多语言 - 基础相当,多语言NanoBEIR(14种语言)也证实了这一趋势。两个模型都能处理长达8192个词元的输入,在长输入检索(NanoLongEmbed)中,a25m是我们比较中最强的密集模型。方法很简单,将22层多语言编码器mmBERT - small剪枝到4/13层,并分两个阶段训练剪枝后的模型——对来自公共语料库的约11亿个多语言对进行大规模对比学习,然后使用8192词元的长文档负样本进行硬负样本微调——使用方向取决于对类型的掩码对比损失以及Matryoshka目标。未使用教师蒸馏,对于a8m,所有训练在单个GPU上约3天内完成。小AP直接带来速度优势:在相同条件下测量的比较模型中,a8m在CPU和GPU上都是最快的——在x86 CPU上比多语言 - e5 - 小快1.6倍,在Raspberry Pi 5上也是最快的。384维输出(可截断为256/128/64)使相似性搜索和索引成本低廉,词汇嵌入的逐行int8量化将a8m的ONNX/OpenVINO版本缩小到124 MiB,可通过此http URL在浏览器中运行。为支持可重复研究,我们发布了模型权重、完整的第一阶段语料库以及独立挖掘的第二阶段硬负样本。

英文摘要

How small can a competitive multilingual retrieval model be? We present Bekko Embedding: its smallest model, bekko-embedding-v1-a8m, has just under 8M Active Parameters (AP) -- the non-embedding parameters that dominate inference compute -- yet on official MMTEB Multilingual v2 Retrieval (nDCG@10) it scores 56.2, above the multilingual-e5 family and BGE-M3 (40x the AP) in our comparison. The higher-quality bekko-embedding-v1-a25m (just under 25M AP) reaches 57.5, on par with gte-multilingual-base, and Multilingual NanoBEIR (14 languages) confirms the trend. Both models handle inputs of up to 8192 tokens, and on long-input retrieval (NanoLongEmbed) a25m is the strongest dense model in our comparison. The recipe is deliberately simple. We prune the 22-layer multilingual encoder mmBERT-small to 4 / 13 layers and train the pruned models in two stages -- large-scale contrastive learning on about 1.1 billion multilingual pairs from our public corpus, followed by hard-negative fine-tuning with 8192-token long-document negatives -- with a masked contrastive loss whose direction depends on pair type, plus the Matryoshka objective. No teacher distillation is used, and all training completes on a single GPU in about 3 days for a8m. Small AP pays off directly in speed: among the compared models measured under identical conditions, a8m is the fastest on both CPU and GPU -- 1.6x multilingual-e5-small on x86 CPU -- and the fastest on a Raspberry Pi 5. The 384-dimensional output (truncatable to 256/128/64) keeps similarity search and indexing cheap, and row-wise int8 quantization of the vocabulary embedding shrinks the a8m ONNX / OpenVINO build to 124 MiB, which runs in the browser via Transformers.js. To support reproducible research, we release the model weights, the complete stage-1 corpus, and the independently mined stage-2 hard negatives.

Comments45 pages, 9 figures. Model weights and training data: https://huggingface.co/hotchpotch

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑