发表机构
Arm, Inc.; Rochester Institute of Technology(Arm公司; 罗切斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出XOR-Trellis,通过超低复杂度反量化器和曲率感知目标,实现无需Hadamard变换的高质量超低位宽LLM量化,兼顾并行重建效率与精度。
AI 中文摘要
Trellis编码量化能够在超低位宽下实现大语言模型(LLM)权重的高维压缩,而无需传统向量量化所需的指数级码本。然而,实际部署面临两个挑战:以足够的并行吞吐量重建压缩权重,以避免反量化成为推理瓶颈;以及在无需昂贵的非相干性变换的情况下保持量化精度。我们通过两种互补技术解决这些挑战。首先,我们引入了一种超低复杂度的trellis反量化器,它使用结构化的、硬件高效的状态到值映射,同时为trellis搜索保留多样化的重建选择。其次,我们重新制定了离散trellis路径优化,采用曲率感知目标,该目标直接在原始坐标空间中反映模型敏感性。这些技术共同实现了高质量的超低位宽trellis量化,具有廉价、高度并行的运行时重建,并且不依赖基于Hadamard的非相干性处理。
英文摘要
Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponentially large codebooks required by conventional vector quantization. Practical deployment, however, presents two challenges: reconstructing compressed weights at sufficient parallel throughput to avoid making dequantization an inference bottleneck, and maintaining quantization accuracy without costly incoherence transformations. We address these challenges with two complementary techniques. First, we introduce an ultra-low-complexity trellis dequantizer that uses a structured, hardware-efficient state-to-value mapping while preserving diverse reconstruction choices for trellis search. Second, we reformulate discrete trellis path optimization with a curvature-aware objective that reflects model sensitivity directly in the original coordinate space. Together, these techniques enable high-quality ultra-low-bit trellis quantization with inexpensive, highly parallel runtime reconstruction and without relying on Hadamard-based incoherence processing.