嵌入器的困境:大型语言模型(LLM)性能更优,但代价是什么?
The Embedder's Dilemma: LLMs Are Better, but at What Cost?
浏览论文内容
中文总结 AI 辅助
该研究对比LLM与嵌入模型在37项任务的性能与成本,发现二者性能接近但LLM成本高、速度慢,建议嵌入模型用于通用任务,LLM用于推理密集型检索。
中文摘要 AI 辅助
是否应将文本嵌入流水线替换为大型语言模型(LLM)?我们通过对六大类共10个LLM与26个嵌入模型(参数规模1.18亿至140亿)在37项任务(涵盖分类、语义文本相似度(STS)、聚类、成对分类及检索)上开展受控的成本感知对比,解答该问题。总体而言,两种范式表现基本持平:最优LLM(Gemini 3.1 Pro)得分77.6,最优嵌入模型得分77.2,差值仅0.4分。二者优势因任务而异:LLM在推理密集型检索中占优,嵌入模型在分类中占优,聚类、STS及成对分类中二者表现相当。达到这种性能均衡的成本很高:LLM的成本是性能相当的嵌入模型的1431倍之多(每次基准测试运行成本为154美元,而嵌入模型仅0.11美元),且测试的开源LLM在相同GPU上处理token的速度比嵌入模型慢2.5至736倍。推理token占LLM推理成本的28%至81%;在我们的 ablation(消融实验)中,降低推理预算可使多数模型的检索质量保持或提升。帕累托前沿包含领先的嵌入模型及1个LLM(Gemini 3.1 Pro)。这些结果支持分工使用:嵌入模型用于相似度、分类及聚类任务,LLM则保留给推理密集型检索任务。我们的代码、数据集及结果可在此httpsURL公开获取。
英文摘要
Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval. In aggregate the two paradigms are effectively tied: the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) differ by 0.4 points. Their strengths differ by task: LLMs lead on reasoning-heavy retrieval, embedding models lead on classification, and the two match on clustering, STS, and pair classification. Reaching that parity is expensive. An LLM costs up to 1,431x more than an embedding model of comparable quality (USD 154 vs. USD 0.11 per benchmark pass), and the open LLMs tested process tokens 2.5 to 736x more slowly on the same GPU. Reasoning tokens account for 28 to 81% of LLM inference cost; lower reasoning budgets preserve or improve retrieval quality for most models in our ablation. The Pareto frontier contains the leading embedding models and one LLM, Gemini 3.1 Pro. These results support a division of labour: use embedding models for similarity, classification, and clustering, and reserve LLMs for reasoning-intensive retrieval. Our code, datasets, and results are publicly available at https://github.com/embeddings-benchmark/embedders-dilemma.
发表机构
- Harvard University(哈佛大学)
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。