arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

什么使问题更难?用自然语言解释问题难度

What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language

Peng Cui, Qiaoyuan Zheng, Rudolf Debelak, Mrinmaya Sachan

arXiv 2610.01627首次发表:更新:

发表机构

ETH Zurich; EPFL(苏黎世联邦理工学院; 洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出数据驱动方法,利用大语言模型生成并验证自然语言假设,解释问题难度差异,实验证明假设可预测且因果有效。

AI 中文摘要

难度是问题最基本的属性之一:它决定了问题能否有效区分不同能力水平的模型。尽管目前已有多种方法可以自动估计或预测难度,但它们仅输出一个描述性的数字,无法解释最初导致问题困难的根本因素。在本工作中,我们提出了一种数据驱动的方法,能够自动生成并验证自然语言假设,以解释为什么一个问题比另一个问题更难。我们首先利用项目反应理论,从大量大语言模型的回答中估计每个项目的难度。然后,我们采样难易问题对比集,并提示大语言模型提出关于差异的候选解释,随后在留出问题上对这些解释进行验证和筛选。在涵盖数学、逻辑和常识推理的三个数据集上的实验结果表明,我们的方法能够产生可解释且具有预测性的假设。这些假设单独使用时,在预测未见问题难度方面与先进的黑盒难度回归器相当或更优;作为附加特征使用时,它们能进一步提升回归器的性能,表明它们发现了现有模型未能捕捉的难度信号。此外,我们证明,根据假设编辑问题可以按预期方向改变其测量难度,这表明所发现的假设是因果有效的难度因素,而非事后描述。因此,我们的方法将纯粹描述性的难度分数转化为可操作的陈述。

英文摘要

Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑