arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23424econ.GNq-fin.EC

错误且更自信:语言模型参加研究生经济学考试的实地实验

Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam

Piyush Akimitsu

AI总结:

研究语言模型在研究生经济学考试中面对误导性信息的表现,通过GERB实验发现误导性信息降低正确答案概率,不同类型模型受影响无显著差异,错误答案有完整解释且模型自信,不对照答案难以发现错误。

AI中文摘要:

一条误导性信息(添加到问题中的无关段落)会使语言模型更频繁地进行错误推理和给出错误答案。然而,模型仍会写出完整解释,且给出的答案与所展示步骤一致。在研究生经济推理基准(GERB)上,通过在有和没有误导性信息的情况下,让38个语言模型回答60个研究生水平的微观经济学问题(每个问题都有详细设置、验证答案和逐步参考解决方案),采用任务内2×2设计进行展示。结果表明,误导性信息使正确答案概率降低12.3个百分点,约为模型正确回答比例的四分之一,38个模型中有37个在此情况下准确性降低。推理能力并不能提供保护。默认推理模型和无推理模式模型之间、开放权重模型和封闭权重模型之间的影响没有显著差异,不过开放权重模型以更低的每个正确答案成本达到了可比的准确性。误导性信息对模型认为简单的问题损害最大。更糟糕的是,误导性信息甚至会使模型认为问题比其无干扰版本更容易,尽管回答错误的频率更高。这些无干扰问题本身就很难,模型平均每五个问题中正确回答不到三个。错误答案仍带有完整解释,与自身推理连贯,且模型保持自信,所以不对照验证答案就无法发现错误。

英文摘要:

A red herring, an irrelevant passage added to a problem, corrupts a language model's reasoning and, through it, its final answer, while the form of the response survives untouched. The benchmark, called the Graduate Economic Reasoning Benchmark (GERB), is sixty graduate-level microeconomics problems, each a detailed setup with a verified final answer and a step-by-step reference solution. Each problem has two versions, one with the red herring and one without, and each of those is asked in two ways, one requesting an explanation and one not. This is a within-subject $2\times2$ factorial experimental design. Thirty-eight language models answer all four versions of every problem. The clean problems (the control group) are already hard, with the models answering under sixty percent correctly on average. The red herring lowers the probability of a correct final answer by 12.3 percentage points, about a quarter of the models' mean accuracy of 0.525. The damage is largest on the problems the model rates as easy. Reasoning ability confers no protection, as the red herring's effect does not differ detectably across models with and without reasoning ability. It does change how the failure looks, since a model with no reasoning mode repeats one wrong answer across waves while a reasoning model wavers. The red herring also leads a model to rate a problem as easier than its clean version, while answering it wrong more often. Although open- and closed-weight models reach the same accuracy, the open-weight models reach it at a substantially lower cost per correct final answer. The form of the response is preserved even as its substance fails. The model still produces an explanation (explanation given), the final answer still follows from the reasoning shown (coherence), and, in the aggregate, it remains the same across waves (consistency).

↑