AI 中文总结
BLADE是一种轻量级框架,通过构建多粒度检查点、学习紧凑探测层子集等,在保持接近基线准确率的同时,减少大语言模型的推理计算与生成令牌量,提升了LLM推理效率。
AI 中文摘要
大语言模型常通过生成长推理链提升任务性能,但计算资源常被冗余的验证与修正浪费。现有基于探测的早退出方法主要检测明确的自我怀疑表述,遗漏了许多更早的终止机会。将检测范围扩展至普通推理边界可提升覆盖度,但也会呈现高度多样的中间状态,其预测信息可能位于不同隐藏层。我们提出用于高效大语言模型推理的边界扩展与层自适应动态出口框架BLADE,这是一个轻量级框架,通过评估生成的前缀是否足以正确回答来动态终止推理。BLADE从句、自我怀疑和段落边界构建多粒度检查点,并通过重复答案补全推导鲁棒的训练标签;进一步学习信息探测层的紧凑子集,而非依赖固定选择或所有层的昂贵表示。推理时,校准后的预测与检查点特定确认规则结合,以平衡响应速度与过早退出风险。在五个基准和两个Qwen3推理模型上的实验表明,BLADE在保持接近基线准确率的同时,使Qwen3-8B的生成令牌减少24.8%,Qwen3-4B减少15.8%;消融研究进一步证实了多样化检查点与自动层选择的益处,为更高效的大语言模型推理提供了有效方法。
英文摘要
Large language models often improve task performance by generating long reasoning traces, but the resulting computation is frequently wasted on redundant verification and revision. Existing probe-based early-exit approaches mainly inspect explicit self-doubt expressions, leaving many earlier termination opportunities undetected. Expanding inspection to ordinary reasoning boundaries improves coverage, but also exposes highly diverse intermediate states whose predictive information may reside in different hidden layers. We present Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning (BLADE), a lightweight framework that dynamically terminates reasoning by estimating whether the generated prefix is sufficient for correct answering. BLADE constructs multi-granular checkpoints from sentence, self-doubt, and paragraph boundaries, and derives robust training labels through repeated answer completions. It further learns a compact subset of informative probe layers instead of relying on fixed choices or expensive representations from all layers. At inference time, calibrated predictions are combined with checkpoint-specific confirmation rules to balance responsiveness and premature-exit risk. Experiments on five benchmarks and two Qwen3 reasoning models show that BLADE preserves near-baseline accuracy while reducing generated tokens by 24.8% on Qwen3-8B and 15.8% on Qwen3-4B. Ablation studies further confirm the benefits of diverse checkpoints and automatic layer selection, demonstrating an effective approach to more efficient LLM reasoning.
Comments8 pages