arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24477cs.CL

数据集稀缺性限制了多语言嵌入模型的稳健评估:以斯拉夫语言为例

Dataset Scarcity Limits Robust Evaluation of Multilingual Embedding Models: A Case Study of Slavic Languages

Ana Gjorgjevikj, Barbara Koroušić Seljak, Tome Eftimov

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多语言嵌入模型评估不均衡问题,提出二维分析框架并应用于斯拉夫语言子集,发现基准测试稀疏性严重,确定了高可迁移性模型,强调需结合证据强度解读结论。

中文摘要 AI 辅助

多语言文本嵌入模型能够在广泛的自然语言处理任务中实现跨语言的知识迁移,但其在高资源、中等资源和低资源语言之间的评估仍存在极大的不均衡性。本文提出了一个专为分析数据集稀缺情况下多语言嵌入基准测试而设计的二维框架,并将其应用于MTEB基准测试的斯拉夫语言子集。该框架区分了任务特定评估与跨任务评估,同时共同分析三个互补方面:(1) 排名稳健性、(2) 模型一致性、(3) 证据强度。在任务特定层面,我们评估模型排名在排名方法和基准测试数据集构成发生变化时的稳定性;在跨任务层面,我们评估模型在某一语言内不同任务间的泛化能力。为量化基准测试结论的可靠性,我们引入了证据强度分数,该分数考虑了数据集的可用性、多样性和稳健性可评估性。我们的分析显示存在严重的基准测试稀疏性,许多斯拉夫语言-任务对依赖于单一数据集或高度相关的基准测试集合,这限制了得出稳健结论的能力。跨任务分析显示存在一小部分高可迁移性模型,最显著的是llama-embed-nemotron-8b、multilingual-e5-large-instruct和Qwen3-Embedding变体,它们在斯拉夫语言和任务中均表现出色。总体而言,结果表明基准测试排名和稳健性结论必须结合其证据强度的特定指标进行解读,并强调基准测试稀缺性是可靠多语言评估的主要障碍。

英文摘要

Multilingual text embedding models enable cross-lingual transfer of knowledge across a wide range of NLP tasks, but their evaluation remains highly uneven across high-, mid- and low-resource languages. In this paper, we propose a two-dimensional framework, specifically tailored for analyzing multilingual embedding benchmarks under dataset scarcity, and apply it on the Slavic-language subset of the MTEB benchmark. The framework distinguishes between task-specific and cross-task evaluation, while jointly analyzing three complementary aspects: (1) ranking robustness, (2) model consistency, and (3) evidence strength. At the task-specific level, we evaluate the stability of model rankings under changes in ranking methodology and benchmark dataset composition. At the cross-task level, we assess the ability of models to generalize across diverse tasks within a language. To quantify the reliability of benchmark conclusions, we introduce an Evidence Strength Score that accounts for dataset availability, diversity, and robustness assessability. Our analysis reveals severe benchmark sparsity, with many Slavic language-task pairs relying on a single dataset or highly correlated benchmark collections, limiting the ability to draw robust conclusions. The cross-task analysis reveals a small group of highly transferable models, most notably llama-embed-nemotron-8b, multilingual-e5-large-instruct, and Qwen3-Embedding variants, that consistently perform well across Slavic languages and tasks. Overall, the results demonstrate that benchmark rankings and robustness conclusions must be interpreted jointly with certain notation of their evidence strength and highlight benchmark scarcity as a major obstacle to trustworthy multilingual evaluation.

发表机构

  • Jožef Stefan Institute(约热夫·斯泰凡研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑