arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型是否理解上下文?一种基于知识图谱的评估框架

Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework

Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Kamal Premaratne, Uthayasanker Thayasivam

arXiv 2609.30484首次发表:更新:

发表机构

University of Miami; University of Moratuwa(迈阿密大学; 莫勒图沃大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于知识图谱的S3KG评估框架,结合结构与语义相似性度量LLM的上下文理解,并在九个基准上实现F1提升达7.6点、AUROC达0.973。

AI 中文摘要

尽管大型语言模型(LLMs)已展现出卓越的语言能力,但其核心仍存在一个深刻的问题:这些模型是真正理解上下文,还是仅仅在空前规模上进行模式匹配?LLM中的上下文理解是指从给定上下文中正确提取相关信息、将其整合为连贯的内部表示,并基于此进行推理以产生事实一致且上下文有根据的回应的能力。然而,传统方法如双语评估替补(BLEU)和困惑度仅衡量表面层面的性能。这揭示了在问答(QA)中的关键缺口,因为回答必须基于上下文而非仅仅是记忆的关联。为填补这一空白,我们提出了一种新颖的基于知识图谱(KG)的评估框架,用于评估LLM在QA中的上下文理解能力。其核心是语义结构相似性(S3KG),一种将结构信号和语义信号结合为单一分数的混合相似性度量。此外,我们开发了一个诊断分析框架,用于在三元组层面识别和分类推理错误,从而实现对模型失败的细粒度分析。综合来看,在九个基准上,S3KG相比最强基线实现了高达+7.6个点的F1提升,AUROC最高达0.973。

英文摘要

While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses. However, traditional methods such as BiLingual Evaluation Understudy (BLEU) and perplexity simply measure surface-level performance. This reveals a critical gap in question answering (QA), where responses must be contextually grounded rather than simply being memorized associations. To fill this void, we propose a novel knowledge graph (KG) based evaluation framework for LLM contextual understanding in QA. Central to this is Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure combining structural and semantic signals into a single score. In addition, a diagnostic analysis framework is developed to identify and categorize reasoning errors at the triplet level, enabling fine-grained analysis of model failures. Together, across nine benchmarks, S3KG achieves F1 gains of up to $+7.6$ points over the strongest baseline and AUROC up to $0.973$.

CommentsAccepted in : AACL-IJCNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑