arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PermuFormer:代数组合学中排列表示的多任务预训练

PermuFormer: Multi-Task Pretraining for Permutation Representation in Algebraic Combinatorics

Henry Kvinge

arXiv 2609.25438首次发表:更新:

发表机构

Pacific Northwest National Laboratory; University of Washington(太平洋西北国家实验室; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PermuFormer,一个在28亿标记多任务语料上预训练的自回归Transformer,用于代数组合学排列任务,微调后优于从零训练模型和通用语言模型,并分析了其内部解码机制。

AI 中文摘要

多样化的预训练已被证明是一种有效的方法,用于学习可复用的、领域感知的表示,为下游任务的微调提供起点。虽然人工智能在数学领域的大部分兴奋点集中在使用前沿推理模型通过语言媒介解决明确定义的问题上,但窄领域的专用模型仍然是人工智能数学生态系统的重要组成部分。与大型语言模型不同,专用模型通常直接训练于数学对象本身(例如,图、数字序列),而不是描述这些对象的文本表示。然而,从零开始训练专用模型的常见做法可能限制其发展领域感知表示的能力,而这种表示能够捕捉数学的多面性。在本文中,我们描述了一种针对代数组合学中排列相关任务的预训练方法。我们引入了PermuFormer,一个自回归Transformer,在包含28亿个标记的多任务、多编码语料库上进行训练。我们表明,PermuFormer是微调预训练期间未见的基本任务和更复杂的研究级任务的有效起点,其性能经常优于从零开始训练的相同架构、基线多层感知机(MLP)以及规模相当且经过微调的通用语言模型。我们还分析了PermuFormer学习解决训练任务的一些内部机制。例如,我们表明,虽然某些任务可以直接从提示的内部表示中线性解码,但其他任务则需要多轮生成后才能解码答案。

英文摘要

Diverse pretraining has been shown to be an effective method for learning reusable, domain-aware representations that provide a starting point for fine-tuning on downstream tasks. While much of the excitement in AI for math has been concentrated in the use of frontier reasoning models to solve well-specified problems through the medium of language, narrow, specialized models remain an important component of the AI for math ecosystem. In contrast to large language models, specialized models are usually trained directly on the mathematical objects themselves (e.g., graphs, sequences of numbers) rather than the textual descriptions that characterize these objects. However, the common practice of training specialists from scratch may limit their ability to develop domain-aware representations that capture the multifaceted nature of mathematics. In this paper, we describe an approach to pretraining for permutation-focused tasks in algebraic combinatorics. We introduce PermuFormer, an autoregressive transformer trained on a 2.8 billion token multi-task, multi-encoding corpus. We show that PermuFormer is an effective starting point for fine-tuning on basic tasks unseen during pretraining and more complex research-level tasks, frequently outperforming the same architecture trained from scratch, baseline MLPs, and a fine-tuned generic language model of comparable size. We also analyze some of the internal mechanisms by which PermuFormer learns to solve training tasks. For example, we show that while some tasks can be linearly decoded directly from the internal representation of the prompt, other tasks require multiple rounds of generation before the answer can be decoded.

Comments33 pages. Comments welcome

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑