arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09004cs.CLcs.LG

大型语言模型中上下文理解的评估

Evaluation of Contextual Understanding in Large Language Models

Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Uthayasanker Thayasivam, Kamal Premaratne

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出基于知识图谱的S3KG混合相似度评估框架,用于衡量大型语言模型在问答中的上下文理解能力,并验证其正确性、忠实性与可解释性。

中文摘要 AI 辅助

大型语言模型(LLMs)在各种自然语言处理任务中展现出令人印象深刻的表现,然而它们展现真正上下文理解的能力仍不确定。传统的评估指标,如困惑度、双语评估替补(BLEU)或表面层面的准确率,无法揭示LLMs在提取、整合和推理上下文信息方面的表现如何——这一差距在问答任务中尤为关键,因为模型必须将回答与基于上下文的知识对齐,而非依赖记忆的关联。我们提出了一种新颖的基于知识图谱的评估框架,引入了知识图谱语义结构相似性(S3KG),这是一种将结构相似性和语义相似性整合为连续评估分数的混合相似度度量,同时配备了一个用于分类推理错误的诊断框架。为验证该流程,我们在一个精选的问答(QA)基准上将S3KG与既有指标进行评估,展示了其在衡量LLM生成回答的正确性、忠实性和可解释性方面的有效性。

英文摘要

Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs extract, integrate, and reason over contextual information--a gap particularly critical in question answering, where models must align responses with contextually grounded knowledge rather than memorized associations. We propose a novel knowledge graph-based evaluation framework introducing Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure integrating structural and semantic similarity into a continuous evaluation score, alongside a diagnostic framework for categorizing reasoning errors. To validate this pipeline, we evaluate S3KG against established metrics on a curated question-answer (QA) benchmark, demonstrating its effectiveness in measuring correctness, faithfulness, and interpretability in LLM-generated responses.

发表机构

  • University of Moratuwa(莫拉图瓦大学)
  • University of Miami(迈阿密大学)

机构由 AI 辅助整理,请以论文原文为准。

↑