arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26067cs.CYcs.AI

易陷阱:为何大语言模型低估由误解驱动的难度

The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty

Amanda La Hadi, Muhammad Johan Alibasa, Guanliang Chen, A. Taufiq Asyhari

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现,用于教育评估难度估计的LLM会系统性低估由学习者误解驱动的数学题难度,提出“易陷阱”现象,其仅能捕捉难度大致排序,无法反映学生实际认知难度。

中文摘要 AI 辅助

大语言模型(LLMs)正越来越多地被用于教育评估中的题目难度估计。然而,这类估计是否反映了学习者实际感受到的难度,目前仍不清楚。本研究调查了LLM生成的难度评分与基础数学任务的学生经验表现之间的一致性。四个广泛使用的基于LLM的系统对32道算术题目进行了1-100分制的难度评分,共生成640个评分(多轮运行),这些评分通过经典测验理论(CTT)和项目反应理论(2PL)与770名印尼本科生的答题得出的经验难度进行比较。结果显示存在中等程度的秩相关(斯皮尔曼ρ=0.52-0.70),表明LLM能够捕捉题目难度的大致排序。但在分数题目中出现了显著且系统性的不一致,一些被LLM持续评为简单的题目对学生而言却是最难的,例如一道正确率仅为34.16%的题目:100: 1/2。我们认为,LLM近似于课程难度,即基于教学顺序应是简单的内容,而非由学习者误解驱动的认知难度,这导致对误解驱动题目的系统性低估,我们将这一现象命名为“易陷阱”。这些发现凸显了基于LLM的难度估计的关键局限,表明在无经验依据的情况下依赖此类估计可能会在评估设计和自适应系统中引入偏差。

英文摘要

Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such estimates reflect how learners actually experience difficulty. This study investigates the alignment between LLM-generated difficulty ratings and empirical student performance on basic mathematics tasks. Four widely used LLM-based systems generated difficulty ratings on a 1-100 scale for 32 arithmetic items across multiple runs (N = 640 ratings). These were compared with empirical difficulty derived from responses of 770 Indonesian undergraduates using Classical Test Theory (CTT) and Item Response Theory (2PL). Results show moderate rank correlations (Spearman's rho = 0.52-0.70), indicating that LLMs capture coarse ordering of item difficulty. However, substantial and systematic misalignment emerges in fraction items. Several items consistently rated as easy by LLMs were among the most difficult for students, such as an item with only 34.16% correct for 100 : 1/2. We argue that LLMs approximate curricular difficulty, or what should be easy based on instructional sequencing, rather than cognitive difficulty driven by learner misconceptions. This leads to systematic underestimation of misconception-driven items, a phenomenon we term the Easy Trap. These findings highlight a critical limitation of LLM-based difficulty estimation and suggest that relying on such estimates without empirical grounding may introduce bias in assessment design and adaptive systems.

发表机构

  • Monash University(莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑