发表机构
Graduate School of Informatics, Nagoya University(名古屋大学信息学研究科)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型语言模型能否复现人类理解时的隐含丰富问题,构建FrameBench基准并实验发现小型模型有挑战、部分大型模型超人类参考分数,还发布了数据集与代码。
AI 中文摘要
在框架语义中,句子理解被认为是通过将词汇意义与称为语义框架的背景知识关联起来进行的,从而使读者能够隐含地用未陈述的信息丰富文本。近期的大型语言模型(LLMs)在广泛的下游任务中取得了优异的性能,但目前尚不清楚它们是否能复现人类在理解过程中自然进行的那种隐含丰富。为解决这一问题,我们引入了基于框架语义的基准FrameBench。FrameBench由多项选择题组成,用于测试模型是否能区分同一动词在不同语境中唤起的框架。我们使用FrameNet式资源以及带有母语者判断的生成-验证流程,构建了英语和日语的基准。我们在多种不同模型上开展的实验揭示了小型模型面临的挑战,而若干大型模型的表现超过了人类参考分数。我们在该https URL发布了构建好的FrameBench数据集以及用于数据集构建和评估的代码。
英文摘要
In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background knowledge called semantic frames, thereby enabling readers to implicitly enrich the text with unstated information. Recent large language models (LLMs) have achieved strong performance across a wide range of downstream tasks. However, it remains unclear whether they can reproduce the kinds of implicit enrichment that humans naturally make during comprehension. To address this question, we introduce FrameBench, a benchmark grounded in frame semantics. FrameBench consists of multiple-choice questions that test whether models distinguish the frames evoked by the same verb across contexts. We construct the benchmark for English and Japanese using FrameNet-style resources and a generation-and-verification pipeline with native-speaker judgments. Our experiments on a diverse set of models reveal challenges for small models, while several large models surpass the human reference scores. We release the constructed FrameBench dataset and the code for dataset construction and evaluation at https://github.com/SasanoLab/FrameBench.
CommentsAccepted in EMNLP Findings 2026