发表机构
Delft University of Technology; Amazon Development Center(代尔夫特理工大学; 亚马逊研发中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有低比特LoRA方法无法直接微调三元Transformer权重的问题,提出三元乘法适配方法,在语言和视觉领域的多个三元模型上恢复了量化损失的性能,且优于相关基线。
AI 中文摘要
三元Transformer具有极高的内存和计算效率,但现有的基于低比特LoRA的方法无法直接微调三元权重。当前方法要么需要反量化,将低比特基础权重恢复到更高精度以与适配权重合并,要么仅更新量化参数,无法得到保持三元属性的合并模型。我们提出三元乘法适配,通过低秩克罗内克分解将三元权重的离散更新(如符号翻转或置零)表示为两个小三元矩阵,逐元素应用于三元权重。该设计参数高效且表达能力强,保留三元域,支持无需反量化的直接合并。在语言和视觉领域的六个模型(包括三元化的LLaMA-3 1B、3B及三元ViT-B/16)上的实验表明,我们的方法可恢复量化损失的大部分性能,且优于强大的低比特和三元基线。代码可在指定URL获取。
英文摘要
Ternary transformers offer extreme memory and compute efficiency, but existing low-bit LoRA-based methods cannot directly fine-tune ternary weights. Current approaches either require dequantization, restoring low-bit base weights to higher precision to merge with adaptation weight, or update only quantization parameters, preventing a merged model that remains ternary. We propose ternary multiplicative adaptation, which represents discrete updates of ternary weights such as sign flips or zeroing through a low-rank Kronecker factorization into two small ternary matrices applied element-wise to ternary weights. This design is parameter-efficient and expressive, preserves the ternary domain, and supports direct merging without dequantization. Experiments on six models across language and vision, including ternarized LLaMA-3 1B and 3B and a ternary ViT-B/16, demonstrate that our method recovers much of the performance lost to quantization and outperforms strong low-bit and ternary baselines. Code is available at https://github.com/alexmanoo/ternary_adaptation.
CommentsAccepted at ECCV 2026. To be published in Volume 17015 of the Lecture Notes in Computer Science series
Journal refComputer Vision - ECCV 2026, Lecture Notes in Computer Science, vol. 17015, pp. 165-181, Springer, 2026
DOI:10.1007/978-3-032-37242-0_10