发表机构
Rutgers University; Harvard University; The Hong Kong University of Science and Technology(罗格斯大学; 哈佛大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出OracleLadder诊断方法,通过递增的里程碑提示定位LLM数学推理失败,发现组合缺口是主要瓶颈,并揭示了不同模型间的性能差异。
AI 中文摘要
大型语言模型(LLM)可以独立解决多步数学问题的每一个中间步骤,但即使给出了步骤路线图及其所有答案,仍然可能无法解决整个问题。我们引入了OracleLadder,一种诊断性评估方法,通过给模型提供递增级别的提示帮助来定位LLM数学推理失败的位置。对于每个问题,一个教师模型会编写一个固定的中间子目标(里程碑)路线图,并由一个确定性的符号验证器对每个答案进行评分。在无帮助、仅有路线图、路线图加里程碑答案以及仅针对每个里程碑单独测试模型,可以将每次失败归类为五种推理缺口之一。在354个NuminaMath问题和六个参数量从8B到671B的模型(Qwen3、gpt-oss、Llama 3.3、DeepSeek-V3.1)上,每个模型的最大缺口都是组合缺口,这是组合性缺口的一种更严格形式。它覆盖了33-48%的问题,在去除LLM审查标记为评分错误的问题后,这一比例为24-37%。准确率和里程碑帮助恢复率对两个最强模型的排名不同,而两次具有相似准确率提升的RLVR运行对问题的移动方式也不同。路线图效应在MATH500和AIME 2024/25上得到复现,在独立的第二教师下,逐问题恢复率在83-87%的问题上一致,帮助阶梯也延续到了代码生成。我们在该https URL发布了数据、路线图、提示和代码。
英文摘要
Large language models (LLMs) can solve every intermediate step of a multi-step math problem on its own and still fail the full problem, even when given a roadmap of the steps and all of their answers. We introduce OracleLadder, a diagnostic evaluation that locates where LLM math reasoning fails by giving the model increasing levels of oracle help. For each problem, a teacher model writes a fixed roadmap of intermediate sub-goals (milestones), and a deterministic symbolic verifier grades every answer. Testing the model with no help, with the roadmap, with the roadmap plus the milestone answers, and on each milestone alone sorts each failure into one of five reasoning gaps. On 354 NuminaMath problems and six models from 8B to 671B parameters (Qwen3, gpt-oss, Llama 3.3, DeepSeek-V3.1), the largest gap for every model is the composition gap, a stricter form of the compositionality gap. It covers 33-48% of problems, and 24-37% after removing problems that an LLM review flags as grading errors. Accuracy and milestone-help recovery rank the two strongest models differently, and two RLVR runs with similar accuracy gains move problems differently. The roadmap effect replicates on MATH500 and AIME 2024/25, per-problem recovery agrees for 83-87% of problems under an independent second teacher, and the help ladder carries over to code generation. We release the data, roadmaps, prompts, and code at https://github.com/slark-prime/OracleLadder.
CommentsAccepted at NeurIPS 2026 (Evaluations and Datasets Track). 47 pages. Code and data: https://github.com/slark-prime/OracleLadder