发表机构
Intel Labs China(英特尔中国实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出ScaleQ-1.58三值后训练量化框架,通过集成AYOT校准方法,提升推理型LLM量化性能,该框架可扩展且泛化性强,仅需少量校准token即可实现优异效果。
AI 中文摘要
我们提出ScaleQ-1.58,这是一种面向推理型大语言模型(LLM)的可扩展三值后训练量化(PTQ)框架。其核心洞见源于一项实证发现:尽管现代LLM通常被训练以展现思维链推理能力,但在PTQ场景下,即便最新的基于学习型可微分三值化的CAT-Q方法,在采用忽略模型推理过程的传统校准方案时,仍会在具有挑战性的数学与编码任务上出现性能崩溃。基于这一发现,我们引入了一种简单的校准方法“关注自身思维(AYOT)”,在三值化过程中,将预训练高精度目标LLM在合适校准样本集上生成的推理轨迹与最终答案,连同对应的问题一起作为上下文输入。ScaleQ-1.58通过将AYOT与CAT-Q简单集成而成,它展现出多项可扩展特性:(1)仅需400万校准token,经ScaleQ-1.58三值化的Qwen3-1.7B,在4项数学与编码任务上的平均性能达到了此前最优BitNet b1.58 2B4T的90.52%以上,而我们的三值化Qwen3-4B则取得了8.97%的绝对性能提升,且量化所需的校准token数量减少了100万倍;(2)ScaleQ-1.58可很好地泛化到密集型与混合专家(MoE)架构,且性能随模型规模增大而提升(最高可达2350亿参数);(3)ScaleQ-1.58在不同难度级别的任务上展现出强泛化能力,涵盖数学、编码、科学逻辑推理,以及常识推理与基础语言生成;(4)其性能随校准token数量增加而持续提升。值得注意的是,AYOT在其他量化比特宽度上也展现出强泛化能力。代码将发布于该https链接。
英文摘要
We propose ScaleQ-1.58, a scalable ternary post-training quantization (PTQ) framework for reasoning LLMs. Its core insight stems from an empirical finding: although modern LLMs are typically trained to exhibit chain-of-thought reasoning capabilities, in the PTQ regime, even the latest CAT-Q method based on learning-based differentiable ternarization still leads to performance collapse on challenging mathematics and coding tasks when using conventional calibration schemes that ignore the model's reasoning process. Driven by this finding, we introduce a simple calibration approach, Attend to Your Own Thoughts (AYOT), where reasoning traces and final answers generated by the pre-trained high-precision target LLM on a proper set of calibration samples are used as the context input during the ternarization process, along with the corresponding questions. ScaleQ-1.58 is formed by simply integrating AYOT with CAT-Q, which demonstrates several scaling properties: (1) with only 4M calibration tokens, Qwen3-1.7B ternarized by ScaleQ-1.58 reaches over 90.52% of the performance of the prior best BitNet b1.58 2B4T averaged over 4 mathematics and coding tasks, and our ternary Qwen3-4B shows an absolute gain of 8.97%, while requiring 1,000,000x fewer calibration tokens for quantization; (2) ScaleQ-1.58 generalizes well to both dense and MoE architectures, with performance improving as model scale increases (up to 235B parameters); (3) ScaleQ-1.58 demonstrates strong generalization across tasks of varying difficulty levels, including mathematics, coding and scientific logic reasoning, as well as commonsense reasoning and basic language generation; (4) its performance continues to improve as the number of calibration tokens increases. Notably, AYOT also exhibits strong generalization ability across other quantization bit-widths. Code will be available at https://github.com/IntelChina-AI/BitTern.
CommentsThis research work was completed and submitted for publication in early May 2026. The project page: https://github.com/IntelChina-AI/BitTern