VALSE:大型语言模型中用于高效推理的垂直自适应层跳过
VALSE: Vertical Adaptive Layer Skipping for Efficient Inference in Large Language Models
AI总结:
本文提出VALSE方法,通过理论证明和轻量级难度估计器实现逐样本非连续层跳过,在保持模型性能的同时提升推理效率。
AI中文摘要:
本文为垂直自适应层跳过建立了理论框架,证明了三个基础性结果:(i)期望FLOPs公式(定理2),给出了任意逐样本跳过调度计算成本的闭式表达式,该表达式是层间跳过概率的函数;(ii)函数空间超集(定理10)和严格包含(定理11)定理,表明跳过层模型严格包含于——但又有意义地逼近——全层函数空间,并给出了一个明确的分离示例;(iii)VALSE与混合专家架构之间的结构对偶性(命题6),将垂直深度稀疏性定位为水平宽度稀疏性的正交对应物。基于这一理论,我们提出了VALSE(垂直自适应层跳过以提高效率),一种逐样本、非连续的层跳过方法:一个轻量级难度估计器从最初的几层对每个输入进行评分,逐层门控选择性地跳过冗余层——包括任意中间层而保留更深层——从而仅为每个输入激活必要的深度,其可行性已在原型规模上进行了初步评估。
英文摘要:
This paper establishes a theoretical framework for vertical adaptive layer skipping, proving three foundational results: (i) an Expected FLOPs formula (theorem 2) giving a closed-form expression for the computational cost of arbitrary per-sample skip schedules as a function of layer-wise skip probabilities; (ii) function-space superset (theorem 10) and strict inclusion (theorem 11) theorems showing that skip-layer models are strictly contained in---yet meaningfully approximate---the full-layer function space, with an explicit separating example; and (iii) a structural duality between VALSE and Mixture-of-Experts architectures (proposition 6), positioning vertical depth-wise sparsity as the orthogonal counterpart to horizontal width-wise sparsity. Building on this theory, we propose VALSE (Vertical Adaptive Layer Skipping for Efficiency), a per-sample, non-contiguous layer skipping method: a lightweight difficulty estimator scores each input from the first few layers, and per-layer gates selectively skip redundant layers---including arbitrary middle layers while retaining deeper ones---so that only the necessary depth is activated for each input, whose feasibility is preliminarily assessed at prototype scale.