开源大语言模型在ESG领域检索增强生成中的实证评估
Empirical Evaluation of Open-Source Large Language Models for Retrieval-Augmented Generation in ESG Domain
- University of Salento(萨伦托大学)
- University of Naples Federico II(那不勒斯费德里科二世大学)
- IFAB Foundation(IFAB基金会)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文基于498份真实ESG报告和100个问答对,评估了七个开源大语言模型在ESG领域RAG系统中的性能,发现检索性能较强但事实正确性偏低,需领域微调,为模型部署提供了数据驱动指导。
AI中文摘要:
环境、社会和治理(ESG)报告对企业问责制至关重要,大语言模型(LLMs)和检索增强生成(RAG)为自动化提取KPI提供了强大潜力。然而,开源LLM在特定领域ESG任务中的性能仍未得到充分理解。本文基于一个结构化框架和评估资源,对ESG情境下的开源LLM进行了评估,该资源基于欧盟上市公司(2010-2024年)的498份真实ESG报告。我们评估了七个开源模型(参数规模从2B到30B)——glm-4.7-flash、nemotron-3-nano:4b、qwen3:4b-instruct、gemma3:4b、gemma4:e4b、gemma4:e2b和ministral-3:8b——使用了100个基于角色的合成问答对,覆盖ESG信息需求。系统性能通过RAGAS指标评估,包括上下文召回率、精确率、相关性、忠实度、答案相关性和事实正确性。结果显示不同架构间性能存在显著差异。各模型的检索性能均较强(上下文召回率约0.58-0.61,上下文精确率约0.78-0.81,上下文相关性0.965-0.985)。生成方面,忠实度差异最大(0.607-0.822),答案相关性差异最小(0.760-0.881):glm-4.7-flash在忠实度上领先(0.822),qwen3在事实正确性上领先(0.449),ministral-3在答案相关性上领先(0.881)。整体较低的事实正确性(0.387-0.449)凸显了领域特定微调的必要性。这项工作为在ESG报告中部署开源模型提供了数据驱动的指导。
英文摘要:
Environmental, Social, and Governance (ESG) reporting is critical for corporate accountability, with Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) offering strong potential to automate KPI extraction. However, open-source LLM performance in domain-specific ESG tasks remains insufficiently understood. This paper evaluates open-source LLMs in ESG contexts using a structured framework and evaluation resource based on 498 real-world ESG reports from EU-listed companies (2010-2024). We evaluate seven open-source models (2B to 30B parameters) -- glm-4.7-flash, nemotron-3-nano:4b, qwen3:4b-instruct, gemma3:4b, gemma4:e4b, gemma4:e2b, and ministral-3:8b -- using 100 persona-based synthetic QA pairs covering ESG information needs. System performance is assessed via RAGAS metrics, including contextual recall, precision, relevance, faithfulness, answer relevancy, and factual correctness. Results show notable performance variations across architectures. Retrieval performance is strong across models (context recall around 0.58-0.61, context precision around 0.78-0.81, context relevance 0.965-0.985). Generation diverges most on faithfulness (0.607-0.822) and least on answer relevancy (0.760-0.881): glm-4.7-flash leads in faithfulness (0.822), qwen3 in factual correctness (0.449), and ministral-3 in answer relevancy (0.881). Low overall factual correctness (0.387-0.449) highlights the need for domain-specific fine-tuning. This work provides data-driven guidance for deploying open-source models in ESG reporting.