arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpecQuant:基于多父量化的推测解码实现自适应LLM推理

SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T

arXiv 2609.21704首次发表:更新:

发表机构

Vellore Institute of Technology(韦洛尔理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SpecQuant是一种无需训练的框架,结合推测解码与多父量化,根据查询复杂度动态选择量化变体,在Qwen2.5模型上实现35-43%加速且精度损失不超过2%,支持设备端高效部署。

AI 中文摘要

在本地运行大型语言模型(LLM)持续受到消费级硬件上计算和内存限制的制约。流行的加速技术,如量化、推测解码和自适应推理,提供了显著的加速效果,但通常需要重新训练、针对每种架构的调优或草稿模型。SpecQuant是一个无需训练的框架,它将推测解码与多父量化相结合,以实现LLM的自适应、高效推理。SpecQuant从共享的基础模型中派生多个量化变体(INT4、FP8、FP16),并根据预测的复杂度动态路由查询;轻量级变体用于简单或事实性任务,而全精度模型用于复杂推理任务或长上下文输入。SpecQuant的共享权重设计确保了推测解码的足够令牌接受率,避免了使用独立草稿父模型时的兼容性问题。我们在基于Qwen2.5的模型上,在MMLU、AlpacaEval和GSM8K数据集或基准上评估了SpecQuant,展示了35-43%的加速,同时精度下降不超过2%,这在LLM社区中是显著的。SpecQuant使得在不同硬件上无需特殊基础设施或专业知识即可实现实用的设备端LLM部署。

英文摘要

Running large language models (LLMs) locally continues to be limited by restrictions of compute and memory on consumer hardware. The popular acceleration technologies, such as quantization, speculative decoding, and adaptive inferencing, offer substantial speed boosts but usually necessitate retraining, per architecture tuning, or draft models. SpecQuant is a trainingfree framework, that combines speculative decoding with multiparent quantization to perform adaptive, efficient inference of LLMs. SpecQuant derives multiple quantized variants (INT4, FP8, FP16) from a shared base model, and dynamically routes queries based on predicted complexity; lightweight variants are used for simple or factual tasks, and full-precision models are used for complex reasoning tasks or long-context inputs. The shared-weight design of SpecQuant ensures sufficient token acceptance for speculative decoding without compatibility issues using separate draft parent models. We evaluate SpecQuant on Qwen2.5 based models on the MMLU, AlpacaEval, and GSM8K datasets, or benchmarks, demonstrating 35-43% speedups without degrading accuracy greater than 2%, substantial within the LLM community. SpecQuant enables practical on-device LLM deployment across diverse hardware without special infrastructure or expertise.

Comments5 pages, 1 figure. Published in the 2026 Fifth International Conference on Power, Control and Computing Technologies (ICPC2T)

Journal ref2026 Fifth International Conference on Power, Control and Computing Technologies (ICPC2T), Raipur, India, 11-13 March 2026, pp. 371-375, IEEE, 2026

DOI:10.1109/ICPC2T68221.2026.11646348

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑