发表机构
Accenture(埃森哲)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过含3个随机种子的严格匹配审计,对比了ConfLayers与SWIFT两种周期步长层跳过方法及训练路由方法的LLM推理效率,发现SWIFT推理速度更快但搜索开销更高,训练路由方法准确率更低,还发布了审计对比模板。
AI 中文摘要
高效大语言模型(LLM)推理的层跳过方法会以一定粒度决定给定输入需执行的Transformer层。我们针对两种周期步长、基于搜索的方法开展了含3个随机种子的严格匹配审计,这两种方法在推理时在线决策,每几步生成步骤重新评估一次:置信门控早停基线ConfLayers,以及真正的自推测解码SWIFT(Xia等人2024),同时结合普通自回归解码,覆盖两个模型规模(Qwen2.5-0.5B和Qwen2.5-1.5B)和两个任务(GSM8K推理与CNN/DailyMail摘要生成)。在四个实验组中,SWIFT在三个组的准确率上表现最强;ConfLayers在所有组均被超越,尤其在1.5B规模的GSM8K任务上差距极大。将在线搜索开销与纯推理成本分离后,SWIFT的真实推理速度在四个组中均快于ConfLayers(快5%-21%),在三个组中反转了原始的挂钟时间排名。ConfLayers的搜索开销小且稳定(占成本的1%-2%),而SWIFT的搜索开销更大且波动(最高达28.7%)。我们还对两种训练路由方法LayerRoute(Sikdar,2026)和LayerDrop(Fan等人2020)进行补充分析,因为它们的决策粒度更粗。在经过验证的协议(含真正的每输入门控、真正的全模型基线、真正的推理时计算跳过)下,两种方法均实现了适度加速(1.08-1.33倍),但准确率远低于周期步长方法,包括1.5B规模的GSM8K任务上LayerRoute几乎完全崩溃(3个随机种子的平均精确匹配率为0.003)。我们发布了完整的审计协议,作为严格匹配的效率对比模板。
英文摘要
Layer-skipping methods for efficient LLM inference decide, at some granularity, which transformer layers to execute for a given input. We present a rigor-matched, three-seed audit of two periodic-step, search-based methods that make this decision online at inference time and re-evaluate it every few generation steps: a confidence-gated early-exit baseline (ConfLayers) and genuine self-speculative decoding (SWIFT, Xia et al. 2024), together with vanilla autoregressive decoding, across Qwen2.5-0.5B and Qwen2.5-1.5B (Yang et al. 2024) on GSM8K (Cobbe et al. 2021) and CNN/DailyMail (Nallapati et al. 2016; See et al. 2017). SWIFT is strongest on accuracy in three of four cells, while ConfLayers is dominated everywhere, with particularly large deficits on GSM8K at 1.5B. Separating online-search overhead from pure inference cost, we find that SWIFT's true inference speed is faster than ConfLayers's in all four cells (5-21%), reversing the naive wall-clock ranking in three; ConfLayers's search overhead is small and stable (1-2% of cost), whereas SWIFT's is larger and more variable across seeds (up to 28.7%). We additionally examine two trained routing methods, LayerRoute (Sikdar, 2026), a per-sequence, input-conditioned hard gate, and LayerDrop (Fan et al. 2020), a fixed, input-independent pruning pattern, as a supplemental analysis rather than a head-to-head comparison because both operate at a coarser decision granularity. Under a verified protocol with genuine per-input gating, a genuine full-model baseline, and genuine inference-time compute skipping, both show modest real speedups (1.08-1.33x) but accuracy well below the periodic-step methods, including near-total LayerRoute collapse on GSM8K at 1.5B (0.003 mean exact-match across three seeds). We release the full audit protocol as a template for rigor-matched efficiency comparisons.
Comments17 pages, 8 figures, 8 tables