量化多步推理的稳定性:通过误差放大
Quantifying the Stability of Multi-Step Reasoning via Error Amplification
浏览论文内容
中文总结 AI 辅助
本文通过误差放大因子量化多步推理稳定性,提出思维链压缩与量化感知训练两种控制方法,在多项任务上平均提升3.5%,长输入提升8.2%。
中文摘要 AI 辅助
我们考虑多步推理过程的稳定性,该过程在语言模型中具有广泛的应用,包括思维链和算法推理。虽然更长的推理序列可以在测试时提高模型的生成能力,但由于中间推理步骤产生的误差在自回归生成中会累积,从而在最终阶段显著增长。在本文中,我们提出以下问题:决定多步推理稳定性的关键因素是什么?首先,我们展示了一个推理误差界,该界限由生成步骤中通过输入空间的雅可比矩阵的谱范数乘积决定。这个乘积可以被视为误差放大因子,它可能随推理步骤数量呈指数增长,作为推理稳定性的定量度量。其次,我们在训练用于预测简单任务(如线性函数和二次函数)的Transformer模型中分析该度量。我们从理论上证明,Transformer模型收敛到一个稳定性度量衰减的解,从而在(任意)长步骤上产生近乎零的推理损失。最后,稳定性分析带来了若干控制稳定性的算法启示,通过(i)思维链长度压缩,降低每一步的敏感性,以及(ii)量化感知训练,正则化输入雅可比范数。我们通过在图形算法推理任务和符号状态跟踪任务上微调语言模型来验证所提出的算法。在七项评估中,我们的算法平均比基线比较提高了3.5%,对于更长长度的输入提高了8.2%。消融分析验证了稳定性度量大幅降低了3-8倍,确认了对(输入空间)雅可比矩阵谱范数的正则化效果。
英文摘要
We consider the stability of multi-step reasoning processes, which have extensive applications in language models, including chain-of-thought and algorithmic reasoning. While longer sequences of reasoning can improve a model's generation capability at test time, the errors due to intermediate reasoning steps can accumulate in autoregressive generation, and thus grow substantially at the end. In this paper, we ask: What are the key factors determining the stability of multi-step reasoning? First, we show an inference error bound governed by the product of spectral norms of the Jacobians taken through the input space across generation steps. This product can be viewed as an error amplification factor, which could scale exponentially with the number of reasoning steps, serving as a quantitative measure of reasoning stability. Second, we analyze this measure in transformer models trained to predict simple tasks like linear and quadratic functions. We theoretically prove that the transformer model converges to a solution where the stability measure decays, thus yielding nearly zero inference loss over (arbitrarily) long steps. Finally, the stability analysis leads to several algorithmic implications for controlling the stability, through (i) chain-of-thought length compression that reduces the sensitivity of each step, and (ii) quantization-aware training that regularizes the input Jacobian norms. We validate the proposed algorithms by fine-tuning language models on graph-algorithmic reasoning tasks and symbolic state-tracking tasks. Across seven evaluations, our algorithms improve over baseline comparisons by 3.5% on average, and by 8.2% for longer-length inputs. Ablation analysis validates that the stability measure is drastically reduced by 3-8$\times$, confirming the regularization effect on the spectral norms of the (input space) Jacobians.
发表机构
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。