arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03401cs.LGcs.AI

更短的推理,更早的答案?对推理接口的评估

Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

Francesca Carlon, Vincent Ginis, Andres Algaba

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估大型语言模型的推理接口,测试提示词和训练设置对推理长度、准确率的影响,发现早答指令和低 effort 推理在特定 token 限制下可提升部分数据集准确率,需综合报告多维度评估指标。

中文摘要 AI 辅助

大型语言模型在回答前往往会进行冗长的推理,这增加了成本和延迟。提示词和训练后的设置可以缩短这种推理,但更短的推理轨迹可能仅表明模型更早停止。在此,我们对198个GPQA Diamond问题和500个MMLU-Pro问题,在匹配的推理跨度下评估同一问题的成对运行结果。我们测试了一个针对Qwen3-14B的数字/简洁提示词,以及gpt-oss-20b和gpt-oss-120b的训练后 effort 设置。Qwen提示词将推理轨迹缩短了12%-17%,而在匹配的 token 限制下,准确率变化很小且参差不齐。简洁/早答指令在512个 token 时将MMLU-Pro的准确率提高了3.8个百分点,在两个运行均未完成时也提高了2.7个百分点;在2048个 token 时,其提升效果不确定。对于gpt-oss,来自已完成的低 effort 和中 effort 推理的候选 logit 答案,比匹配跨度的高 effort 答案准确率高14.5-26.3个百分点。512个 token 的优势大多来自低 effort 更早完成,而未完成运行之间的差异较小且参差不齐。错误的早期答案往往会将概率集中在所选选项上,因此更早停止并不能统一提升概率质量。在这些测试中,严格的截止期限可能倾向于低 effort 或简洁指令,而允许高 effort 完成则可恢复更高的最终准确率。评估应分别报告截止期限前的正确完成情况、运行停止时获得的答案、未完成运行之间的差异以及分配给正确答案的概率。

英文摘要

Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only show that the model stopped sooner. Here, we evaluate paired runs of the same question at matched reasoning horizons across 198 GPQA Diamond and 500 MMLU-Pro questions. We test a numeric/concision prompt that announces a token limit for Qwen3-14B and the trained effort settings of gpt-oss-20b and -120b. The Qwen prompt shortens reasoning traces by 12-17%, while accuracy changes at matched token limits are small and mixed. A concise/early-answer instruction raises MMLU-Pro accuracy by 3.8 percentage points at 512 tokens, including +2.7 points when both runs are unfinished. Its gain at 2,048 tokens is uncertain. For gpt-oss, candidate-logit answers from completed low- and medium-effort reasoning are 14.5-26.3 points more accurate than matched-horizon high-effort answers. Most of the 512-token advantage comes from lower effort finishing earlier, while differences among unfinished runs are smaller and mixed. Wrong early answers often concentrate probability on the chosen option, so earlier stopping does not uniformly improve probability quality. In these tests, a tight deadline can favor lower effort or a concise instruction, whereas allowing high effort to finish can recover higher final accuracy. Evaluations should report correct completion before a deadline, the answer obtained when a run is stopped, differences among unfinished runs, and probability assigned to the correct answer separately.

补充信息

↑