发表机构
School of Mathematics and Statistics, Guizhou University; Institute of Automation, Chinese Academy of Sciences; School of Computer Science and Technology, Guizhou University(贵州大学数学与统计学院; 中国科学院自动化研究所; 贵州大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出用CW复形与锯齿持久同调建模大模型训练中能力变化,证明单胞腔附加仅影响相邻同调类,并区分检查点、过渡与桥敏感类,经多重验证后识别候选能力涌现。
AI 中文摘要
我们提出一个用于大模型训练过程中能力变化的拓扑模型。该模型将知识状态和预注册能力探针表示为有限正则CW复形的子复形。我们证明,附加一个单独的$n$-胞腔只能产生$\nH_n$中的一个类,或消灭$\nH_{n-1}$中的一个类。以可靠性阈值和训练检查点作为两个参数轴,能力的获得与丧失使得沿训练轴的空间非嵌套。并集或交集桥接产生锯齿持久模块,我们区分检查点、过渡和桥敏感类。一个有限四阶段示例通过边界矩阵计算。该框架在选定编码下记录代数变化,并且只有在鲁棒性测试、零模型比较和独立行为验证之后,同调变化才被称为候选能力涌现。
英文摘要
We propose a topological model for capability changes during large-model training. The model represents knowledge states and preregistered capability probes as subcomplexes of a finite regular CW complex. We prove that attaching a single $n$-cell can only create a class in $\mathrm H_n$ or kill a class in $\mathrm H_{n-1}$. With the reliability threshold and the training checkpoint as two parameter axes, capability gains and losses make the spaces along the training axis nonnested. Union or intersection bridges produce zigzag persistence modules, and we distinguish checkpoint, transition, and bridge-sensitive classes. A finite four-stage example is computed by boundary matrices. The framework records algebraic changes under a chosen encoding, and a homological change is called a candidate capability emergence only after robustness tests, null-model comparisons, and independent behavioral validation.
Comments25 pages, 5 figures