arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

选择文本嵌入模型:一个实用的基准测试和决策框架

Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework

Madhav S Baidya

arXiv 2607.23507首次发表:更新:

发表机构

Indian Institute of Technology (BHU) Varanasi(印度理工学院(瓦拉纳西印度理工学院))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对构建检索或搜索系统时文本嵌入模型选择问题,通过基准测试评估 T3EM 等模型,在 MTEB 多任务场景下,追溯从模型到检索结果的全过程,给出依任务、成本等因素选择嵌入模型的实用建议。

AI 中文摘要

选择正确的文本嵌入模型是构建检索或搜索系统中最重要且常被忽视的决策之一。本报告基于基准测试研究开发了一个实用的、基于证据的嵌入模型选择框架。该研究评估了基于商业 API 的嵌入模型 T3EM 与一系列开源替代模型在英语检索任务上的表现,并将结果置于更广泛的大规模文本嵌入基准测试(MTEB)中。报告不仅给出原始基准分数,还追溯了从嵌入模型到检索结果的全过程,包括嵌入如何生成、大规模索引和搜索方式以及文档分块策略如何影响检索质量,从而根据任务、延迟、成本和部署约束为选择嵌入模型提供了一套实用建议。

英文摘要

Choosing the right text embedding model is one of the most consequential -- and most frequently under-examined -- decisions in building a retrieval or search system, yet the model that tops a leaderboard is rarely the best choice for a given deployment. This report develops a practical, evidence-based framework for embedding model selection, built on a benchmarking study that evaluates T3EM (Text 3 Embedding Model), a commercial API-based embedding model, against a broad set of open-source alternatives on English-language retrieval tasks, and situates these findings within the wider Massive Text Embedding Benchmark (MTEB) landscape spanning classification, clustering, semantic similarity, reranking, pair classification, bitext mining, and summarization. Beyond raw benchmark scores, the report traces the full path from embedding model to retrieved result -- how embeddings are produced, how they are indexed and searched at scale, and how document chunking strategy shapes retrieval quality -- so that model choice can be reasoned about as one decision within a complete retrieval pipeline rather than in isolation. The result is a consolidated set of practical recommendations for selecting an embedding model according to task, latency, cost, and deployment constraints.

Comments33 pages, 1 figure, 20 tables. Technical report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑