arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28634cs.CLcs.LG

大语言模型真的能理解题目难度等级吗?对使用大语言模型进行自动题目生成的启示

Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs

Xinyi Wang, Hong Jiao, Ming Li, Sydney Peters, Hanna Choi, Tianyi Zhou, Qingshu Xu

AI总结:

本研究探究LLMs预测题目难度等级的表现,发现GPT-4.1准确率最高但仍低于ConvBERT,LLMs难标注难题,生成特定难度题目需谨慎。

AI中文摘要:

题目难度估计在形成性评估和大规模高风险总结性评估中均发挥关键作用。本研究探究大语言模型(LLMs)在使用大规模读写测试题目预测题目难度等级时的表现,研究了多种提示策略及参数设置,涉及多个LLMs,并将其表现与仅编码器语言模型、基于特征的监督机器学习模型进行比较。零样本设置下温度为0的GPT-4.1取得最高的题目难度等级预测准确率,二次加权Cohen’s kappa(QWK)为0.578;但LLMs的预测准确率低于ConvBERT(QWK=0.625),ConvBERT的表现优于最优的基于特征的监督机器学习模型。进一步分析显示,所有LLMs对难题的标注均存在困难,尤其是当前先进的GPT-5.4倾向于低估题目难度等级。嵌入向量降维结果表明,不同难度等级的题目嵌入向量相互混合,说明仅靠题目自身的语义信息可能不足以预测题目难度等级。研究结果表明,若经验数据显示LLMs无法理解题目难度等级,且随着自身能力提升倾向于将大多数题目视为简单题,则在使用LLMs生成特定难度等级的题目时需谨慎。

英文摘要:

The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. This study explores how large language models (LLMs) perform in predicting item difficulty levels using items from a large-scale Reading and Writing test. The study investigated various prompting strategies and parameter settings across multiple LLMs. LLM performance was compared with encoder-only language models and feature-based supervised machine learning models. Zero-shot GPT-4.1 with a temperature of 0 yielded the highest item difficulty level prediction accuracy, with a quadratic weighted kappa (QWK) of 0.578. However, LLMs' prediction accuracy was lower than that of ConvBERT (QWK = 0.625), which outperformed the best feature-based supervised machine learning model. Further analysis showed that all LLMs struggled to label hard items; in particular, the current advanced GPT-5.4 tended to underestimate item difficulty levels. Dimension reduction of embeddings showed that item embeddings from different difficulty levels were mixed together, indicating that semantic information from items alone is likely insufficient for item difficulty level prediction. The findings suggest that if LLMs cannot understand item difficulty levels as evidenced by empirical data and tend to treat most items as easy when their own capabilities increase, caution should be exercised when using LLMs to generate items with targeted difficulty levels.

补充信息

↑