发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对推理模型对错误数学命题产生谄媚性推导的问题,提出Euston模型,用GraphSynth生成数据微调DeepSeek-R1-8B,将判别差距从-0.5%提升至+27.5%,且不损失数学能力。
AI 中文摘要
推理语言模型被训练为产生解决方案,而非拒绝方案,当交给它们的问题本身是错误的时候,这种偏差依然存在。当被要求证明一个被篡改的定理时,一个强大的模型通常会顺从地生成一个对错误命题的自信推导。我们提出了Euston,一个8B参数的数学声明验证模型,专门训练以抵抗这种行为。训练数据由GraphSynth生成,这是一个概率因子图生成器,它将属性级多样性与时序结构掩码和跨度同步验证相结合,产生了3,026对匹配的真/篡改声明对(共6,052条声明),这些声明来自2010年至2025年间的arXiv论文。我们使用GRPO在基于规则的零API奖励下对DeepSeek-R1-8B进行了189步微调,使用了四块H100 GPU。在一个平衡的200真/200假留出测试集上,平衡准确率从29.50%提升到63.75%,判别差距——即将假声明判定为假的比率与将真声明判定为假的比率之差——从-0.5%(z=-0.1)变为+27.5%(z=+6.0)。关键在于,这一提升并非以牺牲一般数学能力为代价:在官方语义下,AIME 2026的准确率为65.00%,而基线为69.17%,差异为-4.17%,在统计上不显著,而早期在较小的GraphSynth语料库上使用相同配方的运行则崩溃至40.00%。中位响应长度也从19,217个token降至18,296个token,截断率从25.8%降至8.3%,因此改进并非来自更长的思考。我们报告了这一结果以及限制其解释的混杂因素,主要是官方评估集的全假构成以及在现实错误发生率下隐含的低精确率。
英文摘要
Reasoning language models are trained to produce solutions, not to refuse them, and this bias persists when the problem they are handed is false. Asked to prove a corrupted theorem, a strong model will typically comply and produce a confident derivation of something untrue. We present Euston, an 8B mathematical claim-verification model trained to resist exactly this. Training data were generated with GraphSynth, a probabilistic factor-graph generator that couples attribute-level diversity to decode-time structural masking and span-synchronized verification, yielding 3{,}026 matched true/corrupted statement pairs (6,052 statements) drawn from arXiv papers spanning 2010--2025. We fine-tuned DeepSeek-R1-8B with GRPO under a rule-based, zero-API reward for 189 steps on four H100 GPUs. On a balanced 200-true/200-false held-out split, balanced accuracy rises from 29.50% to 63.75% and the discrimination gap---the difference between the rate of calling false statements false and the rate of calling true statements false moves from -0.5% (z=-0.1) to +27.5% (z=+6.0). Critically, the gain is not purchased with general mathematical ability: AIME 2026 accuracy under official semantics is 65.00% against a 69.17% base, a difference of -4.17% that is not statistically significant, whereas an earlier run of the same recipe on a smaller GraphSynth corpus collapsed to 40.00%. Median response length also falls from 19,217 to 18,296 tokens and the truncation rate from 25.8% to 8.3%, so the improvement does not come from thinking longer. We report the result together with the confounds that bound its interpretation, principally the all-false composition of the official evaluation sets and the low precision implied at realistic error prevalence.