数学推理在LLM中按方法而非主题组织
Math Reasoning in LLMs is Organized by Approach, Not Topic
浏览论文内容
中文总结 AI 辅助
本研究通过生成-重放协议和激活签名聚类,证明数学能力LLM的内部计算按推理方法而非主题组织,对基准设计有重要启示。
中文摘要 AI 辅助
数学推理基准通常按主题组织,但语言模型可能转而按可复用的推理方法组织其内部计算。在本文中,我们研究了开放数学能力的LLM在内部是按主题子技能还是按推理方法组织,并提供了证据表明方法是关键。我们引入了一种生成-重放协议:模型首先生成解决方案,之后我们重放精确的提示加生成轨迹,并提取推理标记上的激活重要性签名。我们跨八个模型和五个数学推理来源对这些签名进行无监督聚类,然后通过结构、语义和干预测试评估恢复的结构。在所有40个模型-来源单元中,恢复的聚类优于匹配大小的随机基线。两个独立的前沿LLM评判员发现,在77-82%的真实聚类中具有方法级一致性,而在源内对照中仅为6-11%,且主题纯净的聚类通常获得比主题本身更细粒度的标签。在方法控制的提示中,改变请求的推理方法在八个模型条件中的七个中改变了聚类分配,而释义在很大程度上保持了聚类。这些结果表明,具有数学能力的LLM按推理方法而非基准主题组织内部数学计算。其含义是,按主题分层的基准和主题平衡的训练语料库仍可能错过重要的轴:即使是刻意主题平衡的语料库也可能在推理方法上保持不平衡。
英文摘要
Mathematical reasoning benchmarks are typically organized by topic, but language models may organize their internal computation by reusable reasoning approach instead. In this paper, we investigate whether open math-capable LLMs organize internally by topical sub-skill or by reasoning approach, and we present evidence that the approach is the key. We introduce a generation-replay protocol: a model first generates a solution, after which we replay the exact prompt-plus-generation trajectory and extract activation-importance signatures over the reasoning tokens. We cluster these signatures without supervision across eight models and five mathematical reasoning sources, then evaluate the recovered structure with structural, semantic, and intervention tests. Across all 40 model-source cells, the recovered clusters outperform matched-size random baselines. Two independent frontier-LLM judges find approach-level coherence in 77-82% of real clusters versus 6-11% in within-source controls, and topic-pure clusters usually receive labels finer than the topic itself. In approach-controlled prompting, changing the requested reasoning approach shifts cluster assignment in seven of eight model conditions, whereas paraphrases largely preserve it. These results indicate that math-capable LLMs organize internal mathematical computation by reasoning approach rather than benchmark topic. The implication is that topic-stratified benchmarks and topic-balanced training corpora can still miss the axis that matters: even deliberately topic-balanced corpora may remain imbalanced over reasoning approaches.
发表机构
- Clemson University(克莱姆森大学)
机构由 AI 辅助整理,请以论文原文为准。