发表机构
European University of Tirana(地拉那欧洲大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大型自回归语言模型自我修正盲点问题,提出频谱代数理论SPARC,定义错误传播算子,推导激活阈值,证明基于强化学习的验证器-校正器训练收敛条件,实验验证定理,频谱预测与盲点率匹配。
AI 中文摘要
大型自回归语言模型存在自我修正盲点:当错误归因于外部来源时能可靠修正相同错误,但在自身输出中却无法修正。先前工作通过控制错误注入、错误深度分解、基于强化学习的验证器-校正器训练和内在自我验证记录了这一现象,但缺乏正式模型、定量激活条件和收敛保证。我们用SPARC填补了这些空白,它是自回归生成中自我修正的频谱代数理论。我们定义了错误传播算子,证明了盲点出现的充要条件是该算子的谱半径至少为1。我们推导出校正标记必须超过的尖锐激活阈值,恢复了使用简单“等待”标记观察到的89.3%的盲点减少率。我们还证明了基于强化学习的验证器-校正器训练在验证器-校正器耦合矩阵的谱范数低于1时以与样本数量平方根上的平方耦合强度成比例的速率收敛,且该标准在残差流自回归模态中不变,统一了文本语言模型和自回归图像及视频生成。跨越四个主干和视觉自回归探测器的实验验证了每个定理,频谱预测与测量的盲点率在3.2%的均方根误差内匹配。
英文摘要
Large autoregressive language models exhibit a self-correction blind spot: they reliably fix identical errors when attributed to an external source yet fail to fix the same errors in their own outputs. Prior work has documented this phenomenon empirically, through controlled error injection, error-depth decompositions, RL-based verifier-corrector training, and intrinsic self-verification, but offers no formal model of why generating a token suppresses the ability to detect its error, no quantitative activation condition for correction markers, and no convergence guarantee for reinforcement-learning-based self-correction. We close these gaps with SPARC, a spectral-algebraic theory of self-correction in autoregressive generation. We define the error-propagation operator as the product of per-step attention Jacobians on the residual stream and prove that the blind spot arises if and only if the spectral radius of this operator is at least one. We derive a sharp activation threshold, given as a function of the spectral radius, that a correction marker must exceed, recovering the 89.3\% blind-spot reduction observed with a simple ``Wait'' marker. We further prove that RL-based verifier-corrector training converges at a rate proportional to the squared coupling strength over the square root of the number of samples if and only if the verifier-corrector coupling matrix has spectral norm below one, and that this criterion is invariant across residual-stream autoregressive modalities, unifying text LLMs and autoregressive image and video generation. Experiments across four backbones and a visual autoregressive probe validate every theorem, with spectral predictions matching measured blind-spot rates within 3.2\% RMSE.