arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19916cs.CLcs.AI

KoNeoBench:用于评估LLM对韩国新词理解能力的精选评估数据集

KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms

Soha Lee, Soojin Lee, Heesung Yang, Hyunju Song, Hyunji Lee, Jinsan An, Jeongwan Shin, Jin Hyun Park, Jun Lee, Hyeyoung Park, Kilim Nam

首次发表
浏览论文内容

中文总结 AI 辅助

KoNeoBench是一个基于1,785个韩国新词的精选基准,通过四项任务评估LLM理解韩国新词的能力,实验发现当前模型在源成分恢复、语义区分和定义生成上存在明显局限。

中文摘要 AI 辅助

大型语言模型(LLM)通常在静态基准上进行评估,尽管自然语言通过新出现的词汇和含义不断演变。现有的韩国语基准以既定词汇为中心,因此对此类近期词汇变化的覆盖有限,且其面向英语的设计使得难以评估韩国语的类型学特性,在韩国语中,实义词与功能语素能高效组合。在本文中,我们介绍了KoNeoBench,一个用于评估LLM对韩国新词理解能力的基准。KoNeoBench基于自2020年以来在线新闻中出现的1,785个韩国新词构建,并通过专家词典学审查进行精选。每个条目提供用法示例、构词分析和词典式定义。基于此资源,我们定义了四项任务,并报告了近期模型的结果,同时提供了人类基线。我们的实验表明,当前的LLM在恢复源成分、区分语义类别和生成准确定义方面表现出明显局限性。这些结果揭示了当前LLM仍难以应对的近期韩国词汇变化的具体方面。KoNeoBench可在以下https URL获取。

英文摘要

Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language constantly evolves through newly emerging words and meanings. Existing Korean benchmarks are centered on established vocabulary and therefore provide limited coverage of such recent lexical change, and their English-oriented design makes it difficult to assess the typological properties of Korean, in which content words combine productively with functional morphemes. In this paper, we introduce KoNeoBench, a benchmark for evaluating LLMs' understanding of Korean neologisms. KoNeoBench is built on 1,785 Korean neologisms attested in online news since 2020 and curated through expert lexicographic review. Each entry provides usage examples, word-formation analyses, and dictionary-style definitions. Based on this resource, we define four tasks and report results on recent models, together with a human baseline. Our experiments show that current LLMs exhibit clear limitations in recovering source components, distinguishing semantic categories, and generating accurate definitions. These results reveal specific aspects of recent Korean lexical change that remain challenging for current LLMs. KoNeoBench is available at https://github.com/bcmilab/ko-neobench/ .

发表机构

  • Kyungpook National University(庆北国立大学)
  • Daegu Gyeongbuk Institute of Science and Technology (DGIST)(大邱庆北科学技术院)
  • Texas A&M University(德克萨斯A&M大学)
  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑