ALEE: 通过英语中心的最小对进行嵌入的任意语言评估
ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs
浏览论文内容
中文总结 AI 辅助
提出ALEE框架,利用AMR生成英语最小对并翻译到目标语言,实现任意语言嵌入模型的细粒度诊断,揭示跨语言语义表示中的差距。
中文摘要 AI 辅助
文本嵌入是语义相似性任务的标准方法,但其评估仍然是一个开放的挑战。当前的基准是静态的,仅覆盖有限的语言集,通常具有领域特异性,容易过拟合,且对低资源语言的代表性差。为了解决这些限制,我们引入了ALEE,这是一个将Sentence Smith (Li et al., 2025)扩展到跨语言和段落级别的框架。ALEE使用抽象意义表示(AMR)生成具有受控、细粒度语义偏移的英语最小对,并与目标语言的翻译配对。这种方法使得任何具有英语平行数据的语言模型都能进行有针对性的诊断。我们跨多个嵌入模型和275+种语言(涵盖三个平行数据集)进行了大规模实证研究。在ALEE上,性能在不同语言、文本长度和语言现象之间差异显著,暴露了跨语言语义表示中的持续差距,这些差距与训练资源和子词分词中的语言流行度相关。我们在https://this.url发布ALEE。
英文摘要
Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover only a limited set of languages, are often domain-specific, susceptible to overfitting, and poorly representative of low-resource languages. To address these limitations, we introduce ALEE, a framework that extends Sentence Smith (Li et al., 2025) to the cross-lingual and paragraph level. ALEE uses Abstract Meaning Representations (AMR) to generate English minimal pairs with controlled, fine-grained semantic shifts, which are paired with translations in target languages. This approach enables targeted diagnostics for models in any language with English parallel data. We conduct a large-scale empirical study across a diverse set of embedding models and 275+ languages spanning three parallel datasets. On ALEE, performance varies substantially across languages, text lengths, and linguistic phenomena, exposing persistent gaps in cross-lingual semantic representation that track language prevalence in training resources and subword tokenization. We release ALEE at https://github.com/Andrian0s/any-lang-embed-eval