arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

静态嵌入是否为混合荷兰语检索增加价值?

Do Static Embeddings Add Value to Hybrid Dutch Retrieval?

António Pereira Barata

arXiv 2608.02112首次发表:更新:

AI 中文总结

本研究针对荷兰语检索任务开展受控评估,发现静态嵌入无法为混合检索增加价值,双检索器的词汇-Transformer架构是可靠默认方案。

AI 中文摘要

嵌入基准测试仅衡量模型的独立质量,但无法确定在已结合词汇检索与基于Transformer的检索后,低成本检索器是否能提供互补的排序信息。针对来自荷兰语大规模文本嵌入基准(MTEB-NL)的荷兰语检索任务,我们对该问题开展了受控评估。采用加权倒数秩融合(RRF)方法,将最佳匹配25(BM25)、Qwen/Qwen3-Embedding-0.6B(Qwen)以及两个多语言静态嵌入模型进行融合。对包含14500个查询和786573个文档的五个数据集进行了全面评分,融合权重以0.1为步长在单纯形上进行搜索。采用十折查询级交叉验证,在九折中选择权重并在留出折上评估,通过配对自助法置信区间和符号随机化测试量化所得差异。融合后,荷兰语新闻任务的平均倒数秩(MRR)较训练选出的单一检索器提升0.061,VABB任务提升0.029,WebFAQ NL任务提升0.004,Wikipedia NL任务提升0.025,在Open Tender任务上与BM25表现相当。经Holm校正后,所有四个正向差异仍显著异于零。无任何无限制折为静态检索器分配正向权重:全部50次选择均落在BM25-Qwen边上,强制静态贡献会降低效果。留一数据集选择在每次迭代中均选择相等的BM25-Qwen权重,且在每个留出任务上均优于跨域选择的单一检索器。研究结果表明,在评估的荷兰语任务中,双检索器的词汇-Transformer架构是经测试的可靠默认方案,且独立基准性能不足以确立其在混合检索中的边际价值。

英文摘要

Embedding benchmarks measure standalone model quality, but they do not establish whether a low-cost retriever contributes complementary ranking information once lexical and transformer-based retrieval are already combined. We present a controlled evaluation of this question across Dutch retrieval tasks from the Massive Text Embedding Benchmark for Dutch (MTEB-NL). Weighted reciprocal rank fusion (RRF) combines Best Matching 25 (BM25), Qwen/Qwen3-Embedding-0.6B (Qwen), and two multilingual static embedding models. Five datasets comprising 14,500 queries and 786,573 documents are scored exhaustively, and fusion weights are searched on a simplex in increments of 0.1. Ten-fold query-level cross-validation selects weights on nine folds and evaluates them on the held-out fold; paired bootstrap confidence intervals and sign-randomisation tests quantify the resulting differences. Fusion improves over the training-selected individual retriever by 0.061 mean reciprocal rank (MRR) on Dutch News, 0.029 on VABB, 0.004 on WebFAQ NL, and 0.025 on Wikipedia NL, while matching BM25 on Open Tender. All four positive differences remain distinguishable from zero after Holm correction. No unrestricted fold assigns positive weight to either static retriever: all 50 selections lie on the BM25-Qwen edge, and forcing a static contribution reduces effectiveness. Leave-one-dataset-out selection chooses equal BM25-Qwen weighting in every iteration and outperforms the cross-domain-selected individual retriever on every held-out task. The results support a two-retriever lexical-transformer architecture as a robust tested default across the evaluated Dutch tasks and show that standalone benchmark performance is insufficient to establish marginal value in hybrid retrieval.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑