发表机构
Clemson University; Coastal Carolina University(克莱姆森大学; 卡罗来纳海岸大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现,在微调大语言模型时,按推理方式(如不变量推理、双计数)匹配训练数据比按主题(如概率、数论)匹配更能促进数学迁移,实验显示同方式优于同主题,平均提升10.8至14.3个百分点。
AI 中文摘要
在为大型语言模型选择数学训练数据时,一个自然的组织原则是主题:为概率目标选择概率示例。另一种选择是推理方式:与目标共享解法的完整解答,即使数学领域不同。我们探究在微调后哪种关系能产生更大的迁移效果。我们评估了两个平衡的2×2设计:概率与组合数学分别与不变量推理和双计数交叉(2000个问题),以及数论与几何分别与补集推理和鸽巢原理交叉(800个问题)。在每个设计中,每个单元依次作为留出目标:同方式(SA)来源共享目标的方法但改变主题,而同主题(ST)来源共享主题但改变方法。每个来源在每种角色中出现一次,因此加性来源质量效应在等权聚合对比中相互抵消。在五个基础模型和每个设计三个训练种子下,SA在所有40个种子池化的模型-目标比较中均优于ST。模型级优势在主要设计中从8.2到16.2个百分点(平均10.8),在第二个设计中从12.0到16.0(平均14.3);所有十个模型级95%置信区间均排除零。在两个设计中,ST来源在嵌入和词汇度量下与目标更相似,因此SA优势与语句级相似度的测量排序相反。这些发现表明,在评估的主题-方式组合中,推理方式比主题是更有效的数学迁移匹配标准。
英文摘要
When selecting mathematical training data for LLMs, a natural organizing principle is topic: probability examples for probability targets. An alternative is reasoning approach: worked solutions that share a solution method with the target, even when the mathematical domain differs. We ask which relation produces greater transfer after fine-tuning. We evaluate two counterbalanced $2\times2$ designs: probability and combinatorics crossed with invariant reasoning and double counting (2,000 problems), and number theory and geometry crossed with complement and pigeonhole reasoning (800 problems). In each design, every cell serves as the held-out target in turn: same-approach (SA) sources share the target's method but change the topic, while same-topic (ST) sources share the topic but change the method. Every source appears once in each role, so additive source-quality effects cancel from the equally weighted aggregate contrast. Across five base models and three training seeds per design, SA outperforms ST in all 40 seed-pooled model--target comparisons. Model-level advantages range from 8.2 to 16.2 percentage points in the primary design (mean: 10.8) and from 12.0 to 16.0 in the second design (mean: 14.3); all ten model-level 95% confidence intervals exclude zero. In both designs, ST sources are more similar to targets under embedding and lexical measures, so the SA advantage runs opposite to the measured ordering of statement-level resemblance. These findings identify reasoning approach as a more effective matching criterion than topic for mathematical transfer across the evaluated topic--approach combinations.