arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23440cs.CLcs.AI

推理还是记忆:大语言模型能理解并生成中国歇后语谜语吗?

Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?

Hai Hu, Siyuan Song, Chongtian Shao, Kejia Zhang, Tianjian Zhu, Xiaojing Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

研究通过在新歇后语游戏中测试大语言模型,用多项选择题、自由形式解释生成和新歇后语创作评估其能力,发现前沿中文模型记忆能力较强,新歇后语创作中Gemini 3.1 Pro表现出色,但LLMs整体创造力仍落后于人类专家,凸显数据污染对评估LLMs推理能力的影响。

中文摘要 AI 辅助

在本文中,我们通过在一种汉语游戏——歇后语中测试大语言模型(LLMs)来拓展其推理边界,使用语言学家新创作的歇后语以避免数据污染。我们用多项选择题、自由形式解释生成和新歇后语创作来评估LLMs理解和创作歇后语的能力。在多项选择题中,用现有低频歇后语和新歇后语之间的准确率差值($\Delta_{acc}$)作为记忆指标。以母语者的$\Delta_{acc}$很低,不同模型有不同表现,如前沿中文模型平均$\Delta_{acc}$为23.6%,以英语为中心的模型平均为5.1%。对于新歇后语,Gemini 3.1 Pro表现出色。在歇后语创作方面,LLMs创作的评级远低于人类。这些结果表明考虑数据污染问题,对LLMs推理能力的说法需重新审视,且LLMs在语言相关任务中的创造力仍落后于人类专家。

英文摘要

In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs' ability to understand and create xiehouyu. In MCQ, we use the delta of accuracy ($Δ_{acc}$) between existing but low-frequency xiehouyu and novel ones as an index for memorization. $Δ_{acc}$ for native speakers is very low, suggesting similar processing mechanisms. However, we found that frontier Chinese models have on average a $Δ_{acc}$ of 23.6\%, while English-centric models tested have a mean $Δ_{acc}$ of 5.1\%, suggesting that frontier Chinese models are likely trained with much larger Chinese data, thus memorizing more low-frequency xiehouyu. For novel xiehouyu, Gemini 3.1 Pro demonstrated remarkable ability with acc 92.6, which is 24\% higher than human accuracy. In xiehouyu creation, those created by LLMs receive much worse ratings than those by humans. These results suggest that claims about the reasoning abilities of LLMs may need careful re-examination considering the data contamination issue, and that LLMs' creativity in language-related tasks may still be behind human experts, at least in Chinese xiehouyu.

补充信息

↑