arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型中的分层分级

Hierarchical Grading in Large Language Models

T. Shaska

arXiv 2607.22757首次发表:更新:

发表机构

Oakland University(奥克兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出分级大语言模型(GLLMs)框架,通过代数方法为变换器表示空间分级并传播加权标量作用,扩展相关理论到自回归语言模型。证明分层目标下分级与无分级的极小极大分离,最优分级可离线估计,训练后编译成标准变换器,兼具多种优势。

AI 中文摘要

我们引入了分级大语言模型(GLLMs),这是一个代数框架,它为变换器的表示空间配备了一个分级,并通过嵌入、自注意力和训练目标来传播诱导的加权标量作用。该构建将分级神经网络和分级变换器的理论扩展到自回归语言模型,同时保留了表达能力、渐近计算复杂度和推理成本。分级的好处通过分级环面上的一个Kempf-Ness泛函来体现;比均匀架构有所改进的分级形成一个开放凸锥,其成员资格由一个将分级方向与目标和数据的两个可测量轮廓配对的希尔伯特-芒福德型准则决定;最优分级是两个矩映射的重合点,以封闭形式给出;普通变换器在锥边界上表现为一个半稳定各向同性点:是一个更大分级族的一员而非一个特殊的最优解。另外,对于分层目标,我们证明了分级先验与其不存在之间的极小极大分离:在所有估计器中,分级和均匀目标类的风险在一个明确的样本大小窗口内始终分离,分离因子在几何分层下随层数呈指数衰减。两个轮廓都可以离线估计,所以最优分级解决了一个在训练开始前就得到认证的凸规划。由于分级在训练后被吸收到学习参数中,每个GLLM都编译成具有相同架构和推理复杂度的标准变换器。

英文摘要

We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training objective. The construction extends the theory of graded neural networks and graded transformers to autoregressive language models while preserving expressive power, asymptotic computational complexity, and inference cost. The governing geometric picture is that of geometric invariant theory. The benefit of a grading is expressed by a Kempf--Ness functional on the grading torus; the grades that improve upon the uniform architecture form an open convex cone whose membership is decided by a Hilbert--Mumford-type criterion pairing a grade direction against two measurable profiles of the target and the data; the optimal grades are the coincidence point of two moment maps, given in closed form; and the ordinary transformer appears as a semistable isotropic point on the boundary of the cone: one member of a larger graded family rather than a distinguished optimum. Separately, for level-stratified targets we prove a minimax separation between the graded prior and its absence: over all estimators the risks of the graded and uniform target classes separate throughout an explicit window of sample sizes, by a factor that decays exponentially in the number of levels under geometric stratification. Both profiles are estimable offline, so the optimal grades solve a convex program certified before training begins. Because the grading is absorbed into the learned parameters after training, every GLLM compiles to a standard transformer of identical architecture and inference complexity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑