发表机构
McMaster University; BASF Canada Inc.(麦克马斯特大学; 巴斯夫加拿大公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该工作构建了多语言化学专利CLIR基准,评估八种嵌入模型,发现跨语言检索性能显著下降,强调需更稳健的方法。
AI 中文摘要
跨语言信息检索(CLIR)在跨国行业中日益重要,在这些行业中,关键的技术证据可能以与查询不同的语言存在。然而,现有的基准未能充分捕捉领域特定的跨语言检索,以及聚合召回率所掩盖的检索深度和可恢复性失败。在这项工作中,我们对化学领域的跨语言信息检索进行了基准测试,重点关注专利数据。我们从Google Patents和欧洲专利局(EPO)数据构建了一个多语言数据集,涵盖五种语言(覆盖主要的东方和西方语言),反映了现实世界工业文档的多样性和复杂性。利用该数据集,我们系统地评估了八种最先进的嵌入模型在跨语言检索中的表现。我们的结果显示,单语和跨语言设置之间存在显著的性能差距:对于表现最佳的模型,Recall@10在跨语言设置中从0.72下降到0.53。检索深度也显著下降,在跨语言场景中,相关文档在不同语言中的排名较低。此外,一些在单语设置中表现强劲的多语言嵌入模型,在查询和文档使用不同语言时表现出急剧下降,为跨语言用例中的模型选择提供了实用见解。这些发现突显了当前方法的严重局限性,并强调在领域特定设置中需要更稳健的跨语言检索方法。我们的基准为模型选择提供了可操作的见解,并为工业技术文本上的CLIR建立了一个受控的诊断评估框架。数据和代码可在此https URL公开获取。
英文摘要
Cross-lingual information retrieval (CLIR) is increasingly important in multi-national industries, where critical technical evidence may exist in a different language than the query. However, existing benchmarks do not adequately capture domain-specific cross-lingual retrieval or the retrieval-depth and recoverability failures that aggregate recall hides. In this work, we benchmark CLIR in the chemical domain, with a focus on patent data. We construct a multilingual dataset from Google Patents and the European Patent Office (EPO) data, spanning five languages (covering major Eastern and Western languages) and reflecting the diversity and complexity of real-world industrial documentation. Using this dataset, we systematically evaluate eight state-of-the-art embedding models for cross-lingual retrieval. Our results show a substantial performance gap between monolingual and cross-lingual settings: for the best-performing model, Recall@10 drops from 0.72 to 0.53 in cross-lingual setting. Retrieval depth also degrades significantly, with relevant documents ranked lower across languages in cross-lingual scenarios. Furthermore, some multilingual embedding models that perform strongly in monolingual settings exhibit sharp declines when queries and documents are in different languages, providing practical insights for model selection in cross-lingual use cases. These findings highlight critical limitations of current approaches and emphasize the need for more robust cross-lingual retrieval methods in domain-specific settings. Our benchmark provides actionable insights for model selection and establishes a controlled diagnostic evaluation framework for CLIR over industrial technical text. Data and code are publicly available at https://github.com/MohammadKhodadad/Multi-Lingual-QAC.
CommentsAccepted to the EMNLP 2026 Industry Track. 22 pages including references and appendices; 7 pages of main text