ESG数据嵌入模型基准测试
Benchmarking Embedding Models for ESG Data
浏览论文内容
中文总结 AI 辅助
本研究构建ESG领域基准数据集,系统评估14种开源与闭源嵌入模型在检索及RAG任务中的性能,发现基于Qwen3的模型表现最优,为ESG应用选型提供实践指导。
中文摘要 AI 辅助
环境、社会和治理(ESG)数据的使用对于现代企业问责、可持续发展报告和财务决策至关重要。嵌入模型已成为将非结构化ESG文本转换为适用于下游自然语言处理(NLP)任务的数值表示的有力方法。然而,它们在ESG特定任务中的有效性尚未得到系统研究。在本文中,我们构建了一个专门针对ESG领域的基准数据集。我们对十四个模型进行了基准测试,包括开源和闭源嵌入模型,比较了它们在检索和检索增强生成(RAG)方面的性能。结果表明,不同模型之间存在性能差异,其中基于Qwen3的模型取得了最高的整体性能。本研究为哪些模型更适合ESG RAG任务提供了实用见解。
英文摘要
The use of Environmental, Social, and Governance (ESG) data is fundamental for modern corporate accountability, sustainability reporting, and financial decision-making. Embedding models have emerged as a powerful approach for transforming unstructured ESG text into numerical representations suitable for downstream natural language processing (NLP) tasks. However, their effectiveness in these ESG-specific tasks has not been systematically studied. In this paper, we construct a benchmark dataset specifically tailored to the ESG domain. We benchmark fourteen models, both open-source and closed-source embedding models, comparing their performance with respect to retrieval, and Retrieval-Augmented Generation (RAG). The results demonstrate performance variations across different models, with Qwen3-based models achieving the highest overall performance. This study provides practical insights into which models are better suited for ESG RAG tasks.
发表机构
- University of Salento(萨伦托大学)
- IFAB Foundation(IFAB基金会)
机构由 AI 辅助整理,请以论文原文为准。