QuantForge:发现MXFP4训练后量化的残差分解
QuantForge: Discovering Residual Decompositions for MXFP4 Post-Training Quantization
浏览论文内容
中文总结 AI 辅助
QuantForge通过记录竞争性解释并选择控制项来指导程序修订,发现HiRes量化器,在七项任务中取得最优稳健拟合,并提升可迁移PTQ算法的发现效率。
中文摘要 AI 辅助
四比特训练后量化可以减少大型语言模型的内存需求,但在严格的MXFP4 W4A4条件下保持准确性需要协调多个设计选择。坐标变换会改变块编码误差,进而影响通过网络传播的残差。因此,在搜索之前,有用的算法分解并不完全已知。LLM驱动的程序演化提供了一种探索这些选择的方法,但仅凭性能分数并不能解释接下来应更改哪个设计。我们引入了QuantForge,一个PTQ发现系统,该系统记录竞争性解释,选择区分它们的控制项,并检查后续代码是否实现了由此得出的结论。这种残差编译指导程序修订,同时即使原始解释被拒绝,也保留有用的程序。重新测量修订后的程序会揭示下一个需要解决的误差。该过程发现了HiRes,一种固定的MXFP4量化器,它塑造坐标,细化合法代码分配,并沿注意力路径和MLP路径恢复误差。每个阶段都作用于前一阶段执行后测得的残差。在七项任务中,HiRes实现了最低的七模型稳健拟合(0.09300)和32B下最低的量化Fit-7。在匹配预算的LLM驱动程序演化比较中,每次240次评估调用,QuantForge在八次运行中的六次达到了保留的迁移目标,而文本记忆和反思记忆各为三次,仅分数演化仅为一次,尽管评估的新程序更少。这些结果表明,QuantForge通过将受控证据转化为后续程序更改,提高了可迁移PTQ算法的发现效率。
英文摘要
Four-bit post-training quantization can reduce the memory demands of large language models, but preserving accuracy under strict MXFP4 W4A4 requires coordinating several design choices. Coordinate transforms change block-encoding errors, which in turn affect the residuals propagated through the network. The useful algorithmic decomposition is therefore not fully known before search. LLM-driven program evolution offers a way to explore these choices, but performance scores alone do not explain which design should change next. We introduce QuantForge, a PTQ discovery system that records competing explanations, selects controls that distinguish them, and checks that successor code implements the resulting conclusions. This residual compilation guides program revisions while retaining useful programs even when their original explanations are rejected. Remeasuring the revised program reveals the next error to address. This process discovers HiRes, a fixed MXFP4 quantizer that shapes coordinates, refines legal code assignments, and recovers errors along attention and MLP paths. Each stage acts on residuals measured after the preceding stage has executed. Across seven tasks, HiRes achieves the lowest seven-model Robust Fit (0.09300) and the lowest quantized Fit-7 at 32B. In matched-budget comparisons of LLM-driven program evolution, each with 240 evaluator calls, QuantForge reaches a held-out transfer target in six of eight runs, compared with three each for textual memory and reflection memory, and one for score-only evolution, despite evaluating fewer new programs. These results show that QuantForge improves the discovery of transferable PTQ algorithms by turning controlled evidence into subsequent program changes.
发表机构
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。