arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

绑定三进制平面:将PTQTP约束为统一九级量化器,并为磁盘流式混合专家服务提供持久折叠格式

Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving

Matteo Grella

arXiv 2608.08910首次发表:更新:

发表机构

Anthropic(Anthropic公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将PTQTP约束为统一九级量化器,提出持久折叠格式,在DeepSeek-V4-Flash-0731混合专家模型上实现量化,提升解码速度并减小文件体积,且保真度与基线相当,相关成果开源。

AI 中文摘要

PTQTP将大语言模型(LLM)权重矩阵分解为两个三进制(trit)平面,每组带有两个自由尺度。将这些尺度绑定为3的固定比例,可将该分解简化为单一的统一九级量化器,这是一种已知的平衡三进制恒等式。据我们所知,在撰写本文时,本研究是首个将该恒等式作为约束条件应用于PTQTP求解器内部的工作。随后,这两个三进制平面可无损折叠为一个4位代码平面,我们将其作为持久服务表示:磁盘字节、专家缓存字节和内核输入均为相同的4.0625位/权重块,可通过一次整数点积运算完成处理。对于这种结合了比例为3的九级代码、CPU-SIMD内核、SSD专家流式传输及相同持久字节的方案,我们同样未发现先例。我们将此方法应用于DeepSeek-V4-Flash-0731的路由专家,这是一个284B-A13B规模的混合专家模型,从发布的MXFP4专家权重中一次性量化,并在64GB笔记本电脑上从SSD流式传输专家。以专家无损锚定臂为参考对照,针对4.5位Q4_K基线,在每个固定装置上测量单进程性能:绑定模型在第0步的5个固定装置中匹配官方服务API(Q4_K为4/5),在14个捕获的延续步骤中匹配12个(Q4_K为11/14);在100项MMLU子集上得分86,对比基线为84;解码阶段速度提升6.7%;文件体积缩小9%。在这些小型评估规模下未检测到保真度差异,且两个臂之间的每个固定装置级差异都可追溯到单个测量的接近绑定单元。不过,绑定拟合表现出更高的权重重建误差和更差的困惑度,这是代理指标与参考保真度之间的已测量分离。累积主干三值化阶梯和按位固定的aarch64/x86-64内核完善了本报告。所有代码、格式和评估工件均在fucina推理栈中开源。

英文摘要

PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales. Tying the scales to a fixed ratio of three collapses the decomposition into a single uniform nine-level quantizer, a known balanced-ternary identity. To our knowledge, at the time of writing, this work is the first to impose that identity as a constraint inside PTQTP's solver. The two trit planes then fold losslessly into one 4-bit code plane that we make the persistent serving representation: disk bytes, expert-cache bytes, and kernel input are the same 4.0625-bits/weight blocks, consumed in one integer dot pass. For this conjunction (ratio-3 nine-level code, CPU-SIMD kernels, SSD expert streaming, identical persistent bytes) we likewise found no precedent. We apply this to the routed experts of DeepSeek-V4-Flash-0731, a 284B-A13B mixture-of-experts model, quantizing in one shot from the released MXFP4 expert weights and streaming experts from SSD on a 64 GB laptop. Against a 4.5-bit Q4_K baseline, measured one process per fixture with an expert-lossless anchor arm as reference control, the tied model matches the official serving API on 5/5 fixtures at step 0 (Q4_K: 4/5) and 12/14 captured continuation steps (11/14), scores 86 vs. 84 on a 100-item MMLU subset, decodes 6.7% faster in decode phase, and ships 9% smaller files: no detected fidelity difference at these small evaluation sizes, and every fixture-level difference between the arms traces to a single measured near-tie cell. The tied fit nevertheless shows higher weight-reconstruction error and worse perplexity, a measured dissociation between proxy metrics and reference fidelity. A cumulative trunk-ternarization ladder and bitwise-pinned aarch64/x86-64 kernels complete the report. All code, formats, and evaluation artifacts are open source in the fucina inference stack.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑