arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.04071cs.CLcs.AIcs.LG

超越多语言平均值:MTEB-PT,葡萄牙语句子编码器基准测试

Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders

Lucas Hideki Takeuchi Okamura, Alexandre Alcoforado, Anna Helena Reali Costa

首次发表
浏览论文内容

中文总结 AI 辅助

针对葡萄牙语文本嵌入评估不足,构建MTEB-PT基准,统一评估17个模型。发现葡语性能依赖任务,特定微调可提升,用对比监督和MRL微调模型,发布基准、模型及代码。

中文摘要 AI 辅助

尽管葡萄牙语是世界上使用最广泛的语言之一,但在文本嵌入评估中却未得到充分体现。因此,嵌入模型通常根据英语或多语言指标进行选择,而其在葡萄牙语中的有效性仍不明确。我们提出了MTEB-PT,这是一个从MMTEB的一个子集中构建的葡萄牙语基准,包括14个现有的跨语义文本相似性(STS)、分类、检索和重排数据集。我们使用这个基准在统一协议下评估17个开源和闭源嵌入模型。我们的结果表明,葡萄牙语的性能强烈依赖于任务:多语言排名不能可靠地预测跨任务家族的葡萄牙语特定性能,没有一个单一模型在所有设置中占主导地位,并且具有更强长上下文能力的模型在检索和重排等长输入任务上特别有利。该基准还表明,特定语言的微调仍然可以提高葡萄牙语模型的性能,特别是在与适应数据最匹配的任务类型上。为了检验这种效果,我们用葡萄牙语对比监督和套娃表示学习(MRL)对三个有代表性的骨干模型进行了微调。这些基于基准的基线在STS上取得了最强的收益,这与训练期间使用的主要对称监督一致,同时也提高了检索性能,并在维度截断下保持竞争力。我们发布了MTEB-PT基准、微调后的模型以及训练和评估代码。

英文摘要

Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a result, embedding models are often selected based on English or multilingual metrics, while their effectiveness in Portuguese remains unclear. We present MTEB-PT, a Portuguese benchmark constructed from a subset of MMTEB, comprising 14 existing datasets across Semantic Textual Similarity (STS), classification, retrieval, and reranking. We use this benchmark to evaluate 17 open- and closed-source embedding models under a unified protocol. Our results show that Portuguese performance is strongly task-dependent: multilingual rankings do not reliably predict Portuguese-specific performance across task families, no single model dominates all settings, and models with stronger long-context capacity are particularly advantageous on longer-input tasks such as retrieval and reranking. The benchmark also shows that language-specific fine-tuning still improves model performance in Portuguese, especially on task types that match the adaptation data most closely. To examine this effect, we fine-tune three representative backbone models with Portuguese contrastive supervision and Matryoshka Representation Learning (MRL). These benchmark-informed baselines yield their strongest gains on STS, consistent with the predominantly symmetric supervision used during training, while also improving retrieval and remaining competitive under dimensional truncation. We release the MTEB-PT benchmark, the fine-tuned models, and the training and evaluation code.

发表机构

  • Escola Politécnica, Universidade de São Paulo(圣保罗大学理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑