arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

三元语言模型中的能力分层退化

Capability-Stratified Degradation in Ternary Language Models

Anirudh Malik, M Sparsh Mehra, Poojith Devan

arXiv 2608.28809首次发表:更新:

发表机构

OneBit AI(OneBit AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究将Qwen3.5-0.8B量化为三元模型Cloe,发现其能力非均匀退化,微调后在部分任务保留较高性能,三元转换不适用于通用替代但可作为任务特定模型的紧凑基础。

AI 中文摘要

极端低位推理为构建更小模型和受限部署提供了途径,三元语言模型将权重限制在{-1,0,+1},逼近log₂3≈1.585比特/权重的极限。对于预训练模型,实际问题不仅是权重能否被量化,还在于哪些能力得以保留,以及模型是否仍适用于适配。我们通过使用7240万token的量化感知训练(QAT)将Qwen3.5-0.8B(7.52亿参数)转换为三元权重来探究这一问题,得到的模型Cloe在29个基准测试、表示诊断任务及下游微调中接受评估。证据显示存在非均匀退化:线性探针从全精度教师模型的表示中可恢复43.76%的MMLU答案,但从Cloe中仅恢复26.19%(接近随机水平),表明专业事实信息已丢失;不过Cloe在10个任务上仍保留可测性能,平均为教师模型性能的77.1%。关键的是,微调使Cloe在SST-2上提升至89.8%(为匹配教师模型的95.6%),在XSum上达到79.4%的教师保留率。我们将退化归因于量化诱导的信息丢失,以及有限QAT预算导致的不完全恢复;还指出一个评估陷阱:标准答案字母评分失效(Cloe在98.6%的MMLU问题上输出“A”),需采用延续评分。最终,三元转换不适合作为通用替代方案,但作为特定任务模型的紧凑基础仍具价值。

英文摘要

Extreme low-bit inference offers a route toward smaller models and constrained deployment. Ternary language models restrict weights to $\{-1,0,+1\}$, approaching the limit of $\log_2 3 \approx 1.585$ bits/weight. The practical question for a pretrained model is not simply whether weights can be quantised but which capabilities survive and whether it remains useful for adaptation. We explore this by converting Qwen3.5-0.8B (752M parameters) to ternary weights using 72.4M tokens of quantisation-aware training (QAT). The resulting model, Cloe, is evaluated across 29 benchmarks, representation diagnostics, and downstream fine-tuning. The evidence shows non-uniform degradation. A linear probe recovers 43.76% of MMLU answers from the full-precision teacher's representations but only 26.19% from Cloe (near chance), indicating specialist factual information is lost. However, Cloe retains measurable performance on ten tasks, averaging 77.1% of teacher performance. Crucially, fine-tuning raises Cloe to 89.8% on SST-2 (95.6% of the matched teacher) and reaches 79.4% teacher retention on XSum. We attribute degradation to a combination of quantisation-induced information loss and incomplete recovery due to the limited QAT budget. We also highlight an evaluation pitfall: standard answer-letter scoring failed (Cloe emitted "A" on 98.6% of MMLU questions), necessitating continuation scoring. Ultimately, ternary conversion is unsuitable as a drop-in general replacement yet remains valuable as a compact substrate for task-specific models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑