SABER:基于对抗分支探测的大语言模型推理稳定性感知早停方法
SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing
浏览论文内容
中文总结 AI 辅助
SABER是一种无需训练的大语言模型推理早停框架,通过对抗分支探测平衡推理效率与可靠性,可平均减少30.2%-39.8%的推理token消耗且保持相当准确率。
中文摘要 AI 辅助
大型推理模型(LRM)具备强大的推理能力,但长链推理在中间答案经多推理步骤稳定后会变得低效:额外推理几乎无边际收益,却会产生大量推理开销。现有基于置信度或熵的早停方法难以捕捉推理稳定性,而基于一致性的方法依赖多步轨迹一致性,需顺序评估导致早停延迟。为更好平衡效率与可靠性,我们提出SABER,一种无需训练的稳定性感知早停框架,通过对抗分支探测实现:SABER在中间推理状态周围构建简单有效的语义扰动以形成对抗分支,应用轻量探测估计其可能的最终结果,无需完整轨迹展开;当探测结果在分支间一致时,SABER提前退出,否则继续推理。在多个推理基准和模型架构上的实验显示,SABER平均减少30.2%至39.8%的推理token消耗,同时保持与全长度推理相当的准确率。
英文摘要
Large Reasoning Models (LRMs) achieve strong reasoning capabilities, yet long-chain reasoning becomes inefficient once the intermediate answer stabilizes across reasoning steps: additional reasoning yields little marginal benefit while incurring substantial inference cost. Existing early-exit methods based on confidence or entropy poorly capture reasoning stability, while consistency-based approaches rely on multi-step trajectory agreement, requiring sequential evaluations that delay exit. To better balance efficiency and reliability, we propose SABER, a training-free framework for stability-aware early exit via adversarial branch probing. SABER constructs simple yet effective semantic perturbations around intermediate reasoning states to form adversarial branches, and applies lightweight probing to estimate their likely final outcomes without full trajectory rollouts. When the probed outcomes remain consistent across branches, SABER exits early; otherwise, it continues reasoning. Experiments across multiple reasoning benchmarks and model architectures show that SABER reduces reasoning token consumption by 30.2\%--39.8\% on average while maintaining competitive accuracy with full-length reasoning.
发表机构
- Soochow University(苏州大学)
机构由 AI 辅助整理,请以论文原文为准。