发表机构
Fromthesky Research Labs LLC(弗洛姆斯基研究实验室有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出PLGA注意力机制,其在$G_{LM}=I$时精确包含SDPA,还证明了推理崩溃定理,在测试样本上的分块与顺序评分结果与TruthfulQA度量偏差极小,相关证明核心已通过Lean 4机器验证。
AI 中文摘要
来自幂律解码器表示的大语言模型(PLDR-LLM)及其注意力机制幂律图注意力(PLGA),用通过正张量$A_{LM}$逐元素幂律构建的、由输入生成的双线性算子$G_{LM}$,取代了缩放点积注意力(SDPA)的固定双线性形式。该架构已完全指定,并经过固定参考版本验证;相关声明被标记为定理、条件定理、测量结果或猜想。无条件结论:当$G_{LM}=I$时,PLGA精确包含SDPA;$A_{LM}$和$A_P$严格按元素为正,且$A_{LM}$具有Perron-Frobenius结构;DAG正则化器具有NOTEARS的游走计数形式,正性会阻碍精确无环性;在非共振(标准旋转频率满足该条件)下,交换子准则可识别哪些算子保留相对位置依赖。推理崩溃定理:演绎输出的精确输入不变性会将推理崩溃为具有常数算子的广义SDPA。测量到的不变性:相对波动为$10^{-6}$及以下;扰动界可量化但无法证明缓存推理;组装的代理缺失解码余量。在已发布的检查点上测量到条件三阶段机制(旋转旋转、集中、行映射收缩)。在全局Gram下的分块训练与评分有明确的目标暴露;在测试样本上,分块与顺序评分选择相同答案,且在每个项目上与已发布的TruthfulQA概率质量度量的偏差在$5\times 10^{-5}$以内。自组织临界性作为具有内序参数的现象学框架被引入;未解决的声明成为可证伪的猜想。选定的证明核心在Lean 4中经过机器验证。
英文摘要
The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator $G_{LM}$, built from a positive tensor $A_{LM}$ by elementwise power laws. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture. Unconditionally: PLGA contains SDPA exactly at $G_{LM}=I$; $A_{LM}$ and $A_P$ are strictly entrywise positive, with Perron-Frobenius structure on $A_{LM}$; the DAG regularizer has the NOTEARS walk-counting form and positivity obstructs exact acyclicity; and, under nonresonance (satisfied by standard rotary frequencies), a commutant criterion identifies which operators preserve relative-position dependence. An inference-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator. Measured invariance: relative fluctuations of $10^{-6}$ and below; perturbation bounds quantify but do not certify cached inference; the assembled proxy misses the decoding margin. A conditional three-stage mechanism (rotary twirl, concentration, row-map contraction) is measured on a released checkpoint. Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability-mass metric within $5\times 10^{-5}$ per item. Self-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures. Selected proof cores are machine-checked in Lean 4.
Comments61 pages, 1 figure, 8 tables