测试时数学推理中的预算边界效应
Budget Boundary Effects in Test-Time Mathematical Reasoning
浏览论文内容
中文总结 AI 辅助
本研究通过19,200条轨迹重放,揭示测试时数学推理中严格与建议预算边界的选择差异,并指出预算曲线需联合报告上限、成本、候选、停止规则及选择器信息。
中文摘要 AI 辅助
累积令牌上限可能落在数学推导过程中,迫使测试时控制器在严格模式(在达到上限时停止)和建议模式(允许当前尝试完成)之间做出选择。我们通过配对离线重放19,200条公开轨迹来测量这一边界选择:涉及120道AIME、BrUMO和HMMT问题,以及一个模型的两种存档配置。候选顺序和16次尝试上限固定,答案选择对参考答案和正确性标签不可见。研究发现三点。第一,在4k上限下,建议模式的大部分准确率提升是将弃权(不执行)替换为正确答案;严格模式则为一个未完成的前缀付出代价,而仅完成的选择器无法使用该前缀。第二,沿实际成本的比较与同上限比较不同:在低难度下,建议模式4k的准确率高于严格模式8k,且平均完成成本相当;而在高难度下,其观测准确率比严格模式32k低0.42个百分点,但仅使用其平均令牌的59%。这些总体比较并未确立同等计算优势或准确率等价性。第三,增加候选覆盖并不保证更高的答案准确率:对数概率选择器在覆盖率上升时准确率下降,包括在源级一致性修复之后。在32k上限下,同上限多数投票准确率差异缩小至1.3个百分点以下。预算曲线应同时说明上限、实际成本、合格候选、停止规则和选择器信息。
英文摘要
A cumulative token cap can fall inside a mathematical derivation, forcing a test-time controller to choose between stopping at the cap (strict) and allowing the current attempt to finish (advisory). We measure this boundary choice with paired offline replays of 19,200 public traces: 120 AIME, BrUMO and HMMT problems and two archive configurations of one model. Candidate order and a 16-attempt cap are fixed, and answer selection is blind to reference answers and correctness labels. Three findings emerge. First, at the 4k cap, most advisory accuracy gains replace abstention with a correct answer; strict stopping pays for an unfinished prefix that the completed-only selector cannot use. Second, comparisons along realized cost differ from same-cap comparisons: advisory 4k in low has higher accuracy than strict 8k at comparable mean completion cost, while in high its observed accuracy is 0.42 points below strict 32k using 59% of its mean tokens. These aggregate comparisons do not establish equal-compute superiority or accuracy equivalence. Third, increased candidate coverage does not guarantee higher answer accuracy: a log-probability selector loses accuracy while coverage rises, including after a source-grade consistency repair. Same-cap majority-accuracy differences shrink below 1.3 percentage points at 32k. Budget curves should jointly state the cap, realized cost, eligible candidates, stopping rule and selector information.
发表机构
- Yantong AI(研通人工智能)
机构由 AI 辅助整理,请以论文原文为准。