arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BitNest:用于内存高效LLM推理加速的位嵌套投机解码

BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration

Chence Yang, Ningxi Cheng, Arash Akbari, Qitao Tan, Qingchan Zhu, Ci Zhang, Changdi Yang, Yanzhi Wang, Wei Niu, Jinhui Wang, Jin Lu, Geng Yuan

arXiv 2610.02800首次发表:更新:

发表机构

University of Georgia; Northeastern University; XPeng Motors; The University of Alabama(佐治亚大学; 东北大学; 小鹏汽车; 阿拉巴马大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BitNest提出位嵌套投机解码框架,将低精度草稿嵌入高精度目标,共享权重,实现95.2%接受率和1.48-1.61倍加速,内存高效。

AI 中文摘要

投机解码通过使用轻量级草稿提出多个令牌进行并行验证来加速自回归生成。然而,现有方法通常需要额外的草稿模型或权重表示,在资源受限设备上引入了不可忽视的内存开销。自投机方法减少了这种开销,但仍面临草稿质量、目标质量和存储效率之间的权衡。我们提出了BitNest,一种位嵌套投机解码框架,将低精度草稿直接嵌入到高精度目标表示中。BitNest不是从预定义目标中推导草稿,而是首先构建一个强大的低精度基础,然后通过残差细化恢复高精度目标,使两个模型能够共享单一的物理权重表示。BitNest进一步将此渐进精度设计扩展到KV缓存,用于长上下文推理。在多个7B-8B边缘友好LLM和多样化工作负载中,BitNest实现了平均投机接受率95.2%,同时紧密保持高精度模型质量,并相对于FP16自回归解码提供了1.48-1.61倍的端到端加速。在所有代表性自投机基线支持的LLaMA模型上,BitNest也实现了持续竞争或更高的解码加速。

英文摘要

Speculative decoding accelerates autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification. However, existing methods often require an additional draft model or weight representation, introducing non-negligible memory overhead on resource-constrained devices. Self-speculative approaches reduce this overhead, yet still face trade-offs between draft quality, target quality, and storage efficiency. We propose BitNest, a bit-nested speculative decoding framework that embeds a low-precision draft directly into the higher-precision target representation. Instead of deriving a draft from a predefined target, BitNest first constructs a strong low-precision base and then recovers the higher-precision target through residual refinement, enabling both models to share a single physical weight representation. BitNest further extends this progressive-precision design to the KV cache for long-context inference. Across multiple 7B--8B edge-friendly LLMs and diverse workloads, BitNest achieves an average speculative acceptance rate of 95.2% while closely preserving higher-precision model quality, and delivers 1.48--1.61x end-to-end speedup over FP16 autoregressive decoding. On the LLaMA models supported by all representative self-speculative baselines, BitNest also achieves consistently competitive or higher decoding speedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑