发表机构
OneBit AI(OneBit AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对Qwen3-4B模型采用KOTMS旋转等技术进行后训练三值化,实现存储压缩,虽任务准确率下降但困惑度上升幅度可控,不过压缩未带来推理速度提升。
AI 中文摘要
超低比特大语言模型可减少存储与内存带宽,但标称的“1.58比特”标签无法完整描述存储表示、保留的能力或运行时行为。本研究对指令调优的4B参数模型Qwen开展端到端后训练转换,采用KOTMS旋转、E2M-ATQ三值化及TWLA的GPTQ式误差补偿;实验仅针对权重,激活值保持16位精度,故省略ILA-AMP。我们评估有效比特核算、任务能力保留、困惑度、校准敏感性、检查点组成及部署行为。最终转换中,量化线性权重每权重采用1.641有效比特,覆盖81.62%的模型参数;在10项能力对比中,准确率从64.5%降至54.7%,退化不均:BoolQ保留84.6%经机会校正的教师性能,ARC-Challenge仅保留43.8%;困惑度在WikiText-2从13.639升至18.748,PTB从24.700升至31.992,C4从19.831升至28.966。后续打包操作保留三值平面与缩放因子,将报告的模型大小从8.29 GiB降至3.96 GiB,困惑度基本不变;第三方独立打包尝试存在损失,未纳入主要成果声明;打包后的成果未进行任务准确率或生成吞吐量的端到端基准测试,初步Triton GEMV微基准在某测试形状下比FP16 cuBLAS慢4.6倍,因此本研究不主张仅通过压缩即可实现更快推理。
英文摘要
Ultra-low-bit language models can reduce storage and memory bandwidth, but a nominal "1.58-bit" label does not fully describe the stored representation, retained capability, or runtime behavior. We study an end-to-end post-training conversion of Qwen, an instruction-tuned 4B-parameter model, using KOTMS rotation, E2M-ATQ ternarization, and GPTQ-style error compensation from TWLA. The experiment is weight-only: activations remain at 16-bit precision, so ILA-AMP is omitted. We evaluate effective bit accounting, task capability retention, perplexity, calibration sensitivity, checkpoint composition, and deployment behavior. The final conversion uses 1.641 effective bits per weight for quantized linear weights, with 81.62% of model parameters targeted. Across ten scored capability comparisons, accuracy falls from 64.5% to 54.7%. Degradation is uneven: BoolQ retains 84.6% chance-corrected teacher performance, while ARC-Challenge retains 43.8%. Perplexity rises from 13.639 to 18.748 on WikiText-2, 24.700 to 31.992 on PTB, and 19.831 to 28.966 on C4. A subsequent packing run preserves the ternary planes and scales, reducing reported model size from 8.29 GiB to 3.96 GiB with essentially unchanged perplexity. A separate third-party packing attempt was lossy and is excluded from the primary artifact claim. The packed artifact has not been benchmarked end-to-end for task accuracy or generation throughput. A preliminary Triton GEMV microbenchmark is 4.6x slower than FP16 cuBLAS on one tested shape. We therefore do not claim that compression alone yields faster inference.
CommentsWeight-only post-training ternarization of a 4B-parameter instruction-tuned language model. Activation quantization and end-to-end generation throughput are outside the scope of the primary evaluation