arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

易完成,难选择:探究大型语言模型在ProverbIT基准上的性能

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark

Enrico Mensa, Lorenzo Zane, Calogero Jerik Scozzaro, Matteo Delsanto, Tommaso Milani, Daniele Paolo Radicioni

arXiv 2608.04670首次发表:更新:

发表机构

University of Turin(都灵大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建了意大利谚语基准ProverbIT,评估13个前沿LLMs的谚语处理能力,发现LLMs虽能完成谚语任务,但在无正确答案的选择题中性能骤降,暴露其依赖记忆而非深层语义理解的局限。

AI 中文摘要

大型语言模型(LLMs)已推动计算语言学变革,并在众多自然语言处理任务中取得显著性能,但在理解这些系统如何处理具有文化嵌入性的语言表达方面仍存在显著差距。本文推出ProverbIT,这是一个包含100道选择题的新型意大利基准,旨在评估LLMs完成意大利谚语的能力。我们评估了13个前沿模型,包括大型推理模型(LRMs)和传统LLMs,涉及三项任务:谚语完成、有正确答案的选择题选择、无正确答案的选择题选择。我们的评估揭示了令人惊讶的结果:尽管几乎所有模型都能通过完成任务展现出对谚语的认知,但在转向无正确答案的选择题格式时,性能大幅下降,即便是最先进的推理模型也出现显著退化。通过对两个LRMs的详细思维链分析,我们发现模型存在强烈的选择字面同义词的偏差,且在推理过程中经常提及正确的谚语结尾,却未能成功识别出这些结尾不在给定选项中。这些发现表明,当前LLMs严重依赖记忆模式而非对文化根基表达的更深层次语义理解,凸显了其在比喻语言理解方面推理能力的重要局限。

英文摘要

Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural language processing tasks, yet significant gaps persist in understanding how these systems process culturally embedded linguistic expressions. This paper introduces ProverbIT, a novel Italian benchmark comprising 100 multiple-choice questions designed to evaluate LLMs' ability to complete Italian proverbs. We assess 13 frontier models, including Large Reasoning Models (LRMs) and traditional LLMs, across three tasks: proverb completion, multiple-choice selection with correct answers, and multiple-choice selection without correct answers. Our evaluation reveals surprising results: while nearly all models demonstrate knowledge of the proverbs through successful completion tasks, performance drops dramatically when transitioning to multiple-choice formats without correct answers, with even state-of-the-art reasoning models showing substantial degradation. Through detailed Chain-of-Thought analysis of two LRMs, we uncover that models exhibit a strong bias toward selecting literal synonyms and frequently mention correct proverb endings during reasoning without successfully identifying their absence from the given options. These findings suggest that current LLMs rely heavily on memorized patterns rather than deeper semantic understanding of culturally grounded expressions, highlighting important limitations in their reasoning capabilities for figurative language comprehension.

Journal refProceedings of the Eleventh Italian Conference on Computational Linguistics (CLiC-it 2025), pages 722-734

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑