发表机构
Universidad de Chile; École Polytechnique; Inria; Inria Chile; CENIA(智利大学; 巴黎综合理工学院; 法国国家信息与自动化研究所; 智利国家信息与自动化研究所; 智利国家人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对比了基于知识图谱的增强与检索增强生成在文化相关问答中的表现,提出G-Retriever方法,在LatamQA数据集上将基础LLM错误率降低72%-78%,并实现零样本多语言迁移。
AI 中文摘要
大型语言模型(LLMs)存在长尾缺陷:特定文化的事实,尤其是关于拉丁美洲等代表性不足地区的事实,在预训练语料中出现频率过低,难以被可靠记忆。检索增强生成(RAG)通过将生成过程锚定于外部文本来解决这一问题,但知识图谱(KGs)等结构化替代方案能更严格地控制进入上下文的内容,并在可解释性和可更新性方面具有潜在优势。我们在LatamQA(一个涵盖八个主题类别的文化基础多项选择数据集)上,将Graph-RAG与标准RAG进行基准对比。在我们的主要设置中,图谱通过KGGen(一个最近的开源领域抽取器)从维基百科文章端到端构建,无需人工整理。G-Retriever与RAG性能相当,并将基础LLM的错误率降低了72%(使用标准知识图谱)和78%(使用基准感知变体),随着图谱朝向任务相关内容定向,与RAG的差距进一步缩小。训练得到的投影无需目标语言微调即可零样本迁移到葡萄牙语,表明其具有多语言覆盖能力。
英文摘要
Large language models (LLMs) suffer from a long-tail deficit: culturally specific facts, particularly those concerning underrepresented regions such as Latin America, appear too rarely in pretraining corpora to be reliably memorized. Retrieval-Augmented Generation (RAG) addresses this by grounding generation in external text, but structured alternatives such as Knowledge Graphs (KGs) offer tighter control over what enters the context, along with potential gains in explainability and updatability. We benchmark Graph-RAG against standard RAG on LatamQA, a culturally grounded multiple-choice dataset spanning eight thematic categories. The graphs are built end-to-end from Wikipedia articles with KGGen, a recent open-domain extractor, without manual curation in our main setting. G-Retriever is competitive with RAG and reduces the error of the base LLM by 72\% with a standard KG and 78\% with a benchmark-aware variant, the gap to RAG narrowing further as the graph is oriented toward task-relevant content. The trained projection transfers zero-shot to Portuguese without target-language fine-tuning, indicating multilingual reach.
CommentsAccepted at EMNLP ORACLE workshop 2026. Camera-ready version