发表机构
Tsinghua University; University of Illinois Urbana-Champaign(清华大学; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大型语言模型解数学题易违反约束的问题,提出无训练的两阶段提示协议CFR,经多数据集及多维度实验验证可提升性能,是针对性的测试时干预措施。
AI 中文摘要
大型语言模型能够推导出看似合理的数学对象,却仍可能违反明确要求,例如省略模运算、返回非整数或使用错误的编码答案形式。我们提出约束优先推理(Constraint-First Reasoning, CFR),这是一种无训练的两阶段提示协议:第一阶段提取并总结问题蕴含的约束,第二阶段求解时将中间及最终结果与该总结进行核对。Routed-CFR仅当纯文本正则表达式路由器检测到限制性提示时才激活两阶段协议,否则使用直接思维链(Chain-of-Thought, CoT)。在AIME、CMIMC、BRUMO和AIMO_AMC数据集上,该方法在多个主干模型上较直接CoT实现了性能提升。我们还报告了惯例控制的路由实验、匹配的提示基线、问题级配对测试、解码鲁棒性、约束质量审计、总词元计数以及OlympiadBench评估。这些分析表明,CFR是一种针对性的测试时干预措施,其效益取决于可恢复的约束和第一阶段提取的可靠性,而非作为数学推理的通用替代品。
英文摘要
Large language models can derive a plausible mathematical object yet still violate explicit requirements--for example, by omitting a modular reduction, returning a non-integer, or using the wrong encoded answer form. We introduce Constraint-First Reasoning (CFR), a training-free two-stage prompting protocol: Stage 1 extracts and summarizes constraints entailed by the problem, and Stage 2 solves while checking intermediate and final results against that summary. Routed-CFR activates the two-stage protocol only when a text-only regex router detects restrictive cues; otherwise it uses direct chain-of-thought (CoT). Across AIME, CMIMC, BRUMO, and AIMO_AMC, the method improves direct CoT on multiple backbones. We further report convention-controlled routing experiments, matched prompting baselines, problem-level paired tests, decoding robustness, constraint-quality audits, total-token accounting, and an OlympiadBench evaluation. These analyses position CFR as a targeted test-time intervention whose benefit depends on recoverable constraints and reliable Stage 1 extraction, rather than as a general-purpose replacement for mathematical reasoning.
Comments53 pages, 5 figures, 36 tables