AI 中文总结
Scale-QLoRA通过仅训练原生4位微缩放大语言模型的逐块缩放字段并冻结E2M1代码,实现代码不变适配器合并,在四个模型和任务上精度无损失,还提升了训练效率与任务切换速度。
AI 中文摘要
将LoRA适配器合并到其基础模型中是标准部署实践:它消除了每次前向传播时适配器的开销,并生成一个独立的检查点,可供任何服务栈加载。但对于原生4位微缩放检查点(NVFP4、MXFP4),这一步并非无代价:合并后的权重必须通过量化器回写,这会重新推导检查点的离散E2M1代码平面(约占该工件字节数的90%),导致部署后的工件与一种量化约定绑定,其生命周期中后续任何涉及代码的操作都可能改变它。若采用朴素方法,这一步不仅脆弱,还会使适配性能下降多达39个百分点,因为对于已处于量化网格上的基础模型,重构的最优解就是该基础模型本身。Scale-QLoRA则仅适配原生的逐块缩放字段,在部署网格上训练这些缩放值,并冻结所有E2M1代码。在固定的原生格式、缩放网格、块布局和代码平面下,合并是逐位精确的恒等操作,合并后的工件具有代码不变性。在四个模型和四个任务上,Scale-QLoRA与感知合并的QAT-LoRA均实现了精度无损失,因此不主张二者的精度排序;二者的结构差异在于,QAT-LoRA通过量化器重新推导代码平面,而Scale-QLoRA则精确保留该代码平面。这种差异会影响生命周期代价:最近舍入实现在测量任务上的差异约为1个百分点,更极端的规则不匹配可能使权重空间工件降至约0%,此处将其报告为敏感性边界而非部署频率。保留代码平面还可在训练中移除权重空间直通估计器(在8B密集模型上每步开销为3.9倍),并支持精确回滚、代码平面去重以及仅缩放任务切换速度提升约125倍。
英文摘要
Merging a LoRA adapter into its base model is standard deployment practice: it removes the runtime adapter's per-forward overhead and leaves a single standalone checkpoint any serving stack can load. On a native 4-bit microscaling checkpoint (NVFP4, MXFP4) that step stops being free. The merged weights must be written back through a quantizer, which re-derives the checkpoint's discrete E2M1 code plane (roughly 90% of the artifact's bytes), so the deployed artifact becomes coupled to one quantization convention, and every later code-touching event in its lifecycle can move it. Done naively the step is worse than fragile: it deletes the adaptation, by up to 39 pp, because against an already-on-grid base the reconstruction optimum is that base. Scale-QLoRA instead adapts only the native per-block scale field, trains those scales on the deployment grid, and freezes every E2M1 code. Within a fixed native format, scale grid, block layout and code plane, merging is then a bit-exact identity and the merged artifact is code-invariant. Across four models and four tasks, Scale-QLoRA and merge-aware QAT-LoRA are both accuracy-lossless, so we claim no accuracy ordering between them; they differ structurally, in that QAT-LoRA re-derives the code plane through a quantizer while Scale-QLoRA preserves it exactly. That difference is what the lifecycle prices: nearest-rounding implementations disagree by about a point on the measured task, and more extreme rule mismatches can drive the weight-space artifact to ~0%, which we report as a sensitivity bound rather than a deployment frequency. Preserving the code plane also drops the weight-space straight-through estimator from training (3.9x per step on the dense 8B model) and enables exact rollback, code-plane deduplication, and a ~125x faster scale-only task swap.