大型语言模型在理解现代汉诗的诗意逻辑方面表现良好吗?
Do Large Language Models Perform Well on Comprehending Poetic Logic in Modern Chinese Poetry?
浏览论文内容
中文总结 AI 辅助
该研究针对当前LLMs在现代汉诗诗意逻辑理解评估上的缺口,构建首个基准Peony,评估发现主流LLMs存在相关理解局限,验证了Peony的有效性与必要性。
中文摘要 AI 辅助
大型语言模型(LLMs)在广泛的自然语言处理(NLP)任务中已取得显著进展,但其理解文学文本,尤其是现代汉诗的能力仍未得到充分探索。现代汉诗独特的文学特征要求采用不同的推理形式才能实现有效理解,与传递明确信息的常规文本不同,现代汉诗独特的“诗意逻辑”需要超越表层语义分析的整体推理方法才能被理解。然而,当前的评估范式在很大程度上忽略了这一关键维度。为解决这一缺口,我们提出Peony,这是首个专门用于评估现代汉诗诗意逻辑的基准。我们将诗意逻辑定义为涵盖诗节、诗句和意象三个层级的四项任务,并基于Peony系统评估与分析了六个主流LLMs,在非思考和思考两种配置下对这些模型进行评估。实验结果揭示了当前LLMs在理解现代汉诗诗意逻辑方面的局限性,并验证了Peony的有效性与必要性。我们的数据集和代码将公开提供。
英文摘要
Large Language Models (LLMs) have achieved significant progress across a wide range of natural language processing (NLP) tasks, yet their ability to understand literary texts, particularly modern Chinese poetry, remains largely unexplored. The unique literary characteristics of modern Chinese poetry necessitate a distinct form of reasoning for effective comprehension. Unlike conventional texts that convey clear information, the unique "poetic logic" of modern Chinese poetry requires a holistic reasoning approach that goes beyond superficial semantic analysis to be understood. However, current evaluation paradigms largely ignore this critical dimension. To address this gap, we propose Peony, the first benchmark specifically designed for evaluating the poetic logic of modern Chinese poetry. We define poetic logic as four tasks across three levels, namely stanza, line, and imagery, and systematically evaluate and analyze six mainstream LLMs based on Peony. We evaluate these models under both non-thinking and thinking configurations. The experimental results reveal the limitations of current LLMs in understanding the poetic logic of modern Chinese poetry and validate the effectiveness and necessity of Peony. Our data and code will be available.
发表机构
- Graduate School of Informatics, Kyoto University(京都大学信息学研究科)
- University of Macau(澳门大学)
- College of Computer Science, Inner Mongolia University(内蒙古大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。