arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09240cs.LGcs.AI

将训练后三值化扩展至Qwen3-8B:能力保持、复现、无损打包与打包执行

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

Anirudh Malik, M Sparsh Mehra, Poojith Devan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究将训练后三值化从Qwen3-4B扩展到8B,通过KOTMS旋转、E2M-ATQ和GPTQ补偿,在保持能力的同时实现无损打包与直接执行,验证了规模放大基线。

中文摘要 AI 辅助

超低位宽语言模型有望减少存储和内存流量,但名义上的“1.58位”标签并未指明部署表示或其执行成本。我们研究了将激进的训练后转换流程从Qwen3-4B扩展到Qwen3-8B的规模放大。该转换采用KOTMS旋转、E2M-ATQ自适应三值化和GPTQ风格误差补偿,采用仅权重的A16配置。我们不声称这些算法是新的。我们的贡献在于端到端的规模放大表征:外部复现门、匹配的4B/8B能力分析、跨语料库困惑度、有效位数核算、无损晶格感知打包以及直接打包执行。8B模型在三个语料库上的困惑度比达到1.361倍,其中WikiText-2、C4和PTB的比率分别为1.318倍、1.393倍和1.371倍。在n=500的八个零样本任务上,平均准确率为64.6%,而FP16为72.4%,对应78.5%的机会校正保留率和7.8个百分点的绝对代价。匹配的4B运行保留了69.6%,带来8.9个百分点的8B优势。打包检查点为8.24 GiB,并将记录的困惑度保持到测量精度。直接打包执行在7.35 GiB中达到15.52个token/秒,而初步打包的GEMV仍比FP16 cuBLAS慢。结果是一个经过验证的规模放大基线:模型大小提高了对激进训练后离散化的鲁棒性,实际序列化已针对测量工件解决,直接执行可行,而更广泛的种子、校准分布和内核优化仍有待探索。

英文摘要

Ultra-low-bit language models promise reductions in storage and memory traffic, but a nominal "1.58-bit" label does not specify the deployed representation or its execution cost. We study a scale-up of an aggressive post-training conversion pipeline from Qwen3-4B to Qwen3-8B. The conversion uses KOTMS rotation, E2M-ATQ adaptive ternarisation, and GPTQ-style error compensation in a weight-only A16 configuration. We do not claim these algorithms as new. Our contribution is the end-to-end scale-up characterisation: an external reproduction gate, matched 4B/8B capability analysis, cross-corpus perplexity, effective-bit accounting, lossless lattice-aware packing, and direct packed execution. The 8B model reaches a three-corpus perplexity ratio of 1.361x, with WikiText-2, C4, and PTB ratios of 1.318x, 1.393x, and 1.371x. On eight zero-shot tasks at n = 500, mean accuracy is 64.6% versus 72.4% for FP16, corresponding to 78.5% chance-corrected retention and a 7.8-point absolute cost. The matched 4B run retains 69.6%, yielding an 8.9-point 8B advantage. The packed checkpoint is 8.24 GiB and preserves the recorded perplexity to measurement precision. Direct packed execution reaches 15.52 tokens/s in 7.35 GiB, while a preliminary packed GEMV remains slower than FP16 cuBLAS. The result is a validated scale-up baseline: model size improves robustness to aggressive post-training discretisation, actual serialisation is solved for the measured artefact, and direct execution is feasible, while broader seeds, calibration distributions, and kernel optimisation remain open.

发表机构

  • OneBit AI

机构由 AI 辅助整理,请以论文原文为准。

↑