arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16096cs.IRcs.CL

商业税负:多跳检索基准中的租用与自建盲区

The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks

Luis M. Sanchez, Kosrow Dehnad

首次发表
浏览论文内容

中文总结 AI 辅助

该研究指出多跳检索基准存在商业部署许可与成本信息盲区,发现NVIDIA Nemotron-3-Embed-8B是首个与开源基准性能持平的商业许可嵌入模型,并量化了索引成本差异。

中文摘要 AI 辅助

企业通过检索将语言模型与自身数据相连,而用于对多跳检索系统进行排名的基准,未包含买家在使用已发布数据前需了解的两个关键信息:检索主干是否可商业部署,以及构建成本。在许可方面,该领域的密集检索基准NV-Embed-v2采用cc-by-nc-4.0许可;我们审计的四个领先MuSiQue系统(HippoRAG-2、PropRAG、SAG、KET-RAG)中,三个依赖它来获得最佳结果,且均未披露此情况。在性能方面,我们在统一MuSiQue测试框架中对8家厂商的13种嵌入模型进行了全流程自助抽样置信区间测量。截至2026年年中,存在实际商业税负:最佳商业许可嵌入模型与基准的Recall@5指标相差2.31个点(95%置信区间[0.91,3.71],p=0.001);2026年7月16日发布的NVIDIA Nemotron-3-Embed-8B已缩小差距:Recall@5提升0.24个点(95%置信区间[-0.94,+1.43],p=0.69),Recall@10相差-0.58个点(p=0.28),它与基准持平但未超越,是唯一采用商业许可、可自托管且表现与基准无差异的模型;其他满足前两个条件的模型则比基准低5.2至14.6个点。持久发现是付费与免费的划分:API嵌入模型每次重新索引按token收费,自托管模型免费。在成本方面,审计的5个系统中3个(含微软GraphRAG)未披露索引成本,唯一已发布的GraphRAG美元数据在第三方论文中相差11倍(索引5.64MB语料库一次的成本为2.30美元 vs 24.94美元);推算至1TB时,该未披露选择会导致约42.8万美元至460万美元的成本差异。我们的成本模型将一次性嵌入与 recurring 回答分开:1TB规模下,嵌入成本比图构建成本低7.5至900倍,每天1万次查询的一年回答成本比图构建成本低350倍以上。

英文摘要

Enterprises connect language models to their own data through retrieval. The benchmarks that rank multi-hop retrieval systems leave out two facts a buyer needs before a published number can be used: whether the retrieval backbone may be deployed commercially, and what it costs to build. On licensing: the field's dense-retrieval anchor, NV-Embed-v2, is licensed cc-by-nc-4.0. Of the four leading MuSiQue systems we audit (HippoRAG-2, PropRAG, SAG, KET-RAG), three depend on it for their best numbers and none says so. On performance: we measure thirteen embedders from eight makers on one identical MuSiQue harness with bootstrap confidence intervals throughout. Until mid-2026 there was a real commercial tax: the best commercially-licensed embedder trailed the anchor by 2.31 Recall@5 points (95% CI [0.91, 3.71], p=0.001). NVIDIA's Nemotron-3-Embed-8B, released 2026-07-16, has closed it: +0.24 at Recall@5 (95% CI [-0.94, +1.43], p=0.69), -0.58 at Recall@10 (p=0.28). It matches the anchor, does not beat it, and is the only entrant that is commercially licensed, free to self-host, and indistinguishable from the anchor; every other entrant meeting the first two conditions sits 5.2 to 14.6 points below. The durable finding is the paid-versus-free divide: API embedders charge per token on every re-index, self-hosted ones charge nothing. On cost: three of five audited systems (adding Microsoft's GraphRAG) do not disclose indexing cost, and the only published GraphRAG dollar figures span 11x inside one third-party paper (USD 2.30 vs USD 24.94 to index a 5.64 MB corpus once); extrapolated to 1 TB that undisclosed choice separates roughly USD 428K from $4.6M. Our cost model keeps one-time embedding apart from recurring answering: at 1 TB, embedding sits 7.5x-900x below graph construction, and a year of answering at 10,000 queries/day sits 350x or more below it.

发表机构

  • Columbia University(哥伦比亚大学)
  • New York Institute and Laboratory for Artificial Intelligence(纽约人工智能研究所与实验室)
  • SGX Analytics(SGX分析公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑