发表机构
The Hong Kong University of Science and Technology; Duke Kunshan University(香港科技大学; 昆山杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Loopy提出一种面向循环语言模型的低比特量化框架,通过循环深度感知的目标函数和渐进式校准窗口分配,在目标部署深度选择最优共享低比特表示,显著降低困惑度并提升性能。
AI 中文摘要
循环语言模型通过重复执行共享的循环核心,提供了一种参数高效的方式来扩展迭代式的测试时计算。训练后量化(PTQ)可以减少循环语言模型的内存占用和推理成本,但量化共享核心引入的误差会影响后续的核心。在PTQ方法中,通道缩放和正交旋转保留了浮点计算,同时产生具有不同量化质量的表示。我们发现量化配置候选的排名会随循环深度而变化,这促使我们在目标部署深度进行配置选择。然而,在该深度下对整个校准集评估每个候选者的成本很高。因此,我们提出了Loopy,一种PTQ框架,通过循环深度感知的目标函数来形式化共享核心量化,根据目标部署深度下的最终预测损失来选择共享的低比特表示。通道缩放和正交旋转对候选表示进行参数化。为了高效地近似解决这一选择问题,Loopy在保留完整目标深度执行的同时,仅使用前向评估,逐步将校准窗口分配给有希望的候选者。在八种设置中,Loopy在不同基线中取得了最先进的结果。在Ouro-1.4B上,W4A4条件下,Loopy相对于SpinQuant将LAMBADA困惑度降低了36.5%。我们的代码可在以下网址获取:https://this https URL。
英文摘要
Looped language models provide a parameter-efficient way to scale iterative test-time computation by repeatedly executing a shared recurrent core. Post-training quantization (PTQ) can reduce the memory footprint and inference cost of looped language models, but errors introduced by a quantized shared core affect subsequent cores. Among PTQ methods, channel scaling and orthogonal rotations preserve the floating-point computation while producing representations with different quantization quality. We find that quantization configuration candidate rankings can change with recurrent depth, motivating configuration selection at the target deployment depth. However, evaluating every candidate over the full calibration set at this depth is costly. We therefore propose Loopy, a PTQ framework that formulates shared-core quantization through a recurrent-depth-aware objective, selecting shared low-bit representations by their final prediction loss at the target deployment depth. Channel scaling and orthogonal rotations parameterize the candidate representations. To approximately solve this selection problem efficiently, Loopy progressively allocates calibration windows to promising candidates while preserving complete target-depth execution, using only forward evaluations. Across eight settings, Loopy achieves the state-of-the-art results among different baselines. On Ouro-1.4B under W4A4, Loopy reduces LAMBADA perplexity by 36.5% relative to SpinQuant. Our code is available at https://github.com/Shameless0817/Loopy-review.git.
Comments31 pages, 17 figures, 9 tables