AI 中文总结
该研究发现,随着LLM性能提升,标准CoT提示的干扰作用凸显,推理专用模型在零样本设置下的数学推理性能优于少样本CoT提示。
AI 中文摘要
思维链(CoT)提示仍是评估模型推理能力的标准基线,该技术最初用于引导大型语言模型(LLM)输出逐步推理过程,否则模型倾向于直接给出最终答案。但许多现代LLM在面对推理任务时会原生生成CoT风格的响应,这促使我们重新审视标准CoT提示的有效性。我们在数学问题解决任务上评估了多款现代中型语言模型,发现专门用于推理的模型在简单零样本设置下的性能优于使用少样本CoT示例的情况——例如Mathstral在GSM8K数据集上的准确率从约77%提升至约84%,且无需额外成本。对于测试的通用型模型,零样本CoT提示也足以优于少样本CoT基线。我们将此归因于“引导-干扰”权衡:标准CoT提示还要求风格适配、格式合规性及潜在的不必要上下文,这会干扰模型聚焦核心推理任务。我们的发现表明,随着模型性能增强,使用标准CoT提示正逐渐成为一种干扰源。
英文摘要
Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was introduced to elicit step-by-step reasoning from large language models (LLMs), which would otherwise tend to directly output the final answer. However, many modern LLMs produce CoT-style responses \textit{natively} when presented with reasoning tasks, which made us revisit the effectiveness of standard CoT prompting. We evaluate several modern mid-sized language models on a math problem-solving task and find that models specialized for reasoning achieve better performance in a simple zero-shot setting than when using few-shot CoT examples - significantly surpassing officially reported results at no additional cost (e.g., from $\sim$77\% to $\sim$84\% for Mathstral on GSM8K). For the tested general-purpose model, a zero-shot CoT prompt is also sufficient to outperform a few-shot CoT baseline. We attribute this to a `guidance-distraction' tradeoff: standard CoT prompting also demands style adaptation, formatting compliance, and potentially undesired contextualization, which can distract models from the core reasoning task. Our findings suggest that using standard CoT prompting increasingly acts as a source of distraction as models grow stronger.
Comments10 pages, 3 tables