arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向大语言模型的安全增强的基于种子的权重量化

Security-Enhanced Seed-Based Weight Quantization for Large Language Models

Qiuyu Ren, Sudipta Paria, Aritra Dasgupta, Swarup Bhunia

arXiv 2609.38477首次发表:更新:

发表机构

University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Seed-Q,一种安全增强的敏感性感知的基于种子的权重量化框架,通过非均匀比特分配和LFSR生成权重,在减少比特的同时保持性能,并增强对比特翻转攻击的鲁棒性。

AI 中文摘要

大型语言模型(LLM)会产生大量的存储、内存带宽和能源成本,这促使人们寻求紧凑的权重表示。现有的基于种子的压缩方法从紧凑的伪随机表示中重建权重,但并未明确考虑模型权重的非均匀敏感性。我们引入了Seed-Q,一种安全增强的、敏感性感知的基于种子的权重压缩框架,该框架使用基于轻量级线性反馈移位寄存器(LFSR)的权重生成和非均匀比特分配。我们的方法为敏感权重分配更大的表示预算,同时积极压缩不太敏感的区域。重要的是,这种非均匀分配不需要侧信息:解码器确定性地重建比特分配方案,没有任何层级依赖于解码后的权重,从而无需存储每块元数据或使用校准数据,同时保持基线编码率。在多种不同LLM上的实验表明,Seed-Q以更少的比特达到了SeedLM的4比特困惑度,而在相同的4比特/权重下,与SeedLM相比,它同时减少了困惑度下降和零样本准确率损失。我们还表明,Seed-Q同时实现了对模型参数比特翻转攻击的高安全性,因为比特损坏会影响多个重建的权重,极大地放大了其影响并使其更容易被检测。我们进一步在基于ASIC的加速器中实现了Seed-Q,并展示了与先前基于种子的方法相比适度的硬件开销。

英文摘要

Large language models (LLMs) incur substantial storage, memory-bandwidth and energy costs, motivating compact weight representations. Existing seed-based compression methods reconstruct weights from compact pseudo-random representations but do not explicitly account for the non-uniform sensitivity of model weights. We introduce Seed-Q, a security-enhanced sensitivity-aware seed-based weight compression framework that uses lightweight Linear Feedback Shift Register (LFSR)-based weight generation with non-uniform bit allocation. Our approach assigns larger representation budgets to sensitive weights while aggressively compressing less sensitive regions. Importantly, this non-uniform allocation requires no side-information: the decoder deterministically reconstructs the bit-allocation schedule, with no rung depending on the decoded weights, eliminating the need to store per-block metadata or use calibration data while preserving the baseline coding rate. Experiments across diverse LLMs show that Seed-Q matches 4-bit perplexity of SeedLM with fewer bits, while at the same 4 bits/weight it reduces both perplexity degradation and zero-shot accuracy loss relative to SeedLM. We also show that Seed-Q simultaneously achieves high security against bit-flip attacks on model parameters, as bit corruption affects multiple reconstructed weights, greatly amplifying its impact and making it easier to detect. We further implement Seed-Q in an ASIC-based accelerator and demonstrate modest hardware overhead compared to prior seed-based approaches.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑