arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07853cs.LGcs.CL

迷失在 bf16 转换中:导出三元语言模型可能使大多数低学习率代码更改失效

Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes

Avichal Sahai, Nishant Raj, Animesh Srivastava

首次发表
浏览论文内容

中文总结 AI 辅助

本文审计三元语言模型导出流程,发现 bf16 转换导致大量低学习率微调代码失效并严重降低准确率,提出两种兼容性补救措施达到非劣效标准。

中文摘要 AI 辅助

诸如 BitNet b1.58、Falcon-E 和 BitCPM 之类的三元语言模型使用更高精度的潜在权重进行微调,并作为由导出步骤生成的三元代码进行部署,而在实验室记录的处理流程中,该步骤首先将潜在权重转换为 bf16。我们对三个实验室的处理流程进行了审计。在已发布的检查点中,对已发布潜在权重的 fp32 量化与部署代码在 Falcon-E 和 BitCPM 中 0.83% 至 1.77% 的代码上不一致,在 BitNet 2B-4T 中不一致率为 1.530%;对于 Falcon-E 和 BitCPM,大多数不一致是 bf16 舍入恰好落在阈值上,而 ties-to-even 规则将其映射为零,且未经修改的 onebitllms 导出器逐字节复现了所有四个 Falcon-E 发布版本。在微调端点,选择学习率以匹配标称的学习率与 bf16 最小单位(ULP)之比时,文档记录的导出将 Falcon-E-1B-Base 的贪婪 GSM8K 严格准确率从 58.79% 降至 0.78%,将 BitCPM-CANN-0.5B 的从 36.13% 降至 0.39%,而 bf16 保存和重新加载使 BitNet 2B-4T 的严格准确率下降 27.54 个百分点,同时其最后数字准确率上升。两种兼容性补救措施——直接写入训练量化器的代码,或调整 bf16 输入直到未修改的工具输出这些代码——在所有三个模型中均达到了针对在线评估的 4 个百分点严格准确率非劣效标准。在两个模型系列中,对初始阈值距离的随机干预支持了微调所改变的代码的依赖于距离的选择。

英文摘要

Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts the latents to bf16. We audit those pipelines across three labs. In released checkpoints, fp32 quantization of the shipped latents disagrees with the deployed codes on 0.83-1.77% of codes in Falcon-E and BitCPM and on 1.530% in BitNet 2B-4T; for Falcon-E and BitCPM most disagreements are products that bf16 rounding lands exactly on the threshold, which ties-to-even maps to zero, and the unmodified onebitllms exporter reproduces all four Falcon-E releases byte for byte. At fine-tuned endpoints, with learning rates selected to match a nominal learning-rate-to-bf16-ULP ratio, the documented export lowers greedy GSM8K strict accuracy from 58.79% to 0.78% for Falcon-E-1B-Base and from 36.13% to 0.39% for BitCPM-CANN-0.5B, and a bf16 save and reload lowers BitNet 2B-4T's strict accuracy by 27.54 points while its last-number accuracy rises. Two compatibility remedies, writing the training quantizer's codes directly or adjusting the bf16 inputs until the unchanged tools emit them, each met a 4-point strict-accuracy non-inferiority criterion against online evaluation in all three models. In two model families, randomized interventions on the initial distance from the threshold support distance-dependent selection of the codes that fine-tuning changes.

发表机构

  • Ofbusiness

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑