简单性悖论:揭穿关于大语言模型评估中提示和数据集的神话
ReasonLab: A Controlled and Auditable Evaluation of Prompting Techniques for Multiple-Choice QA
浏览论文内容
中文总结 AI 辅助
研究大语言模型评估中提示和数据集相关问题,通过对8种提示技术在10个MCQA数据集上的实证研究,发现基线提示常优于复杂技术,还研究了相关关键现象,表明LLM评估社区或使提示工程过复杂,存在性能差距,为模型改进提供机会。
中文摘要 AI 辅助
探究大语言模型(LLMs)的能力并为多项选择题问答(MCQA)构建强大解决方案仍是自然语言理解中的核心挑战。此外,LLMs的迅速扩散产生了一种隐含假设,即更复杂的提示技术能带来更好的性能。一些研究声称更复杂的提示技术性能更好,但未提供全面评估。我们通过对8种提示技术在10个MCQA数据集上进行全面实证研究来填补这一空白,涵盖27种模型配置,约4300个独特问题被评估超430000次。我们的发现揭示了一个惊人的悖论:在各种基准测试中,基线提示始终优于复杂推理技术。只有最小专家和归纳角色框架(CoT-Expert和CoT-Inductive)比基线有小但统计上显著约3个百分点的提升,而我们测试的其他复杂技术与之匹配或表现更差,差距常很大(Self-Analogical高达31个百分点)。我们进一步研究了三个关键现象:(1)Qwen3-30B-A3B-Thinking-2507在Elo评级中的意外胜利;(2)不同思维预算的模型变体之间的性能-效率权衡,揭示了依赖模型的最优配置;()数据集难度的巨大差异,60%的基准准确率低于70%,从最容易到最难有47.5个百分点的差距,表明模型有很大改进空间。这些结果表明,LLM评估社区可能使提示工程过于复杂,不同基准之间仍存在巨大性能差距,为真正的模型改进而非提示优化提供了机会。
英文摘要
Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding. Furthermore, the rapid proliferation of LLMs has created the implicit assumption that more sophisticated prompting techniques yield better performance. Several studies claim such gains, but report them under differing models, prompt wordings and answer-extraction rules, so the gains cannot be attributed to the technique alone. We address this gap with ReasonLab, an evaluation framework in which the prompting technique is a first-class experimental variable alongside the model and the dataset, and which retains every generation for inspection. Using ReasonLab we conduct a controlled study of 8 prompting techniques across 10 MCQA datasets, 27 model configurations and 480,927 evaluations at temperature 0. We find that the prompting technique is a minor determinant of accuracy: on configurations without a reasoning budget the reasoning triggers improve on direct prompting by only 3.92 to 4.69 pp and are indistinguishable from one another, and on configurations with reasoning enabled no technique differs by more than 0.51 pp. Self-Generate is the only technique with a consistent effect, a reduction of 2.95 pp. We further investigate three phenomena: (1) the comparison of models on a common set of datasets, where model size does not predict accuracy, (2) the trade-offs across thinking budgets, where enabling reasoning is worth up to 12.74 pp whereas an eightfold budget increase adds only 0.48 to 2.10 pp, and (3) the variation in dataset difficulty, with 60% of benchmarks below 70% accuracy and a 43.9 pp spread from easiest to hardest. These results suggest that, for MCQA, the prompting technique is a minor lever compared with enabling model reasoning, and that substantial headroom remains.