arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LightRot:用于精确低比特大语言模型推理的轻量级旋转方案与架构

LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference

Sangjin Kim, Yuseon Choi, Jungjun Oh, Byeongcheol Kim, Hoi-Jun Yoo

arXiv 2607.27704首次发表:更新:

发表机构

PIM Semiconductor Design Research Center(PIM半导体设计研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出LightRot轻量级旋转方案与硬件加速器,通过算法创新结合28nm工艺实现4比特推理27.4 TOPS/W能效,适配LLaMA系列模型,为低比特LLM推理提供新范式。

AI 中文摘要

随着大语言模型(LLMs)在各领域展现出卓越能力,实现高能效且精确的推理这一挑战愈发关键。本研究提出LightRot,一种专为低比特LLM推理设计的轻量级旋转方案与专用硬件加速器。所提架构将分组局部旋转(GLR)、离群值方向对齐(ODA)算法与基于分层快速哈达玛变换(FHT)的旋转单元相集成,以解决低比特量化中的关键挑战,包括旋转操作的能量开销问题。该加速器采用28nm CMOS工艺实现,在4比特推理下达到27.4 TOPS/W的峰值能效,超越了现有最先进设计。与依赖更高精度推理或在GPT-2等基础语言建模任务上评估的传统方法不同,LightRot针对LLaMA2-13B、LLaMA3-8B等高级模型进行了优化,其性能在MT-Bench上得到进一步验证,展现出对现实世界对话场景的强适用性,并重新定义了基于聊天的AI系统的基准。通过算法创新与硬件效率的协同,本研究为可扩展的低比特LLM推理树立了新范式,为可持续AI发展铺平道路。

英文摘要

As large language models (LLMs) continue to demonstrate exceptional capabilities across various domains, the challenge of achieving energy-efficient and accurate inference becomes increasingly critical. This work presents LightRot, a lightweight rotation scheme and dedicated hardware accelerator designed for low-bit LLM inference. The proposed architecture integrates Grouped Local Rotation (GLR) and Outlier Direction Aligning (ODA) algorithms with a hierarchical Fast Hadamard Transform (FHT)-based rotation unit to address key challenges in low-bit quantization, including the energy overhead of rotation operations. The proposed accelerator, implemented in a 28nm CMOS process, achieves a peak energy efficiency of 27.4 TOPS/W for 4-bit inference, surpassing prior state-of-the-art designs. Unlike conventional approaches that rely on higher-precision inference or evaluate on basic language modeling tasks like GPT-2, LightRot is optimized for advanced models such as LLaMA2-13B and LLaMA3-8B. Its performance is further validated on MT-Bench, demonstrating robust applicability to real-world conversational scenarios and redefining benchmarks for chat-based AI systems. By synergizing algorithmic innovations and hardware efficiency, this work sets a new paradigm for scalable, low-bit LLM inference, paving the way for sustainable AI advancements.

Comments13 pages, journal version. Published in IEEE Journal on Emerging and Selected Topics in Circuits and Systems (JETCAS), vol. 15, no. 2, pp. 231-243, 2025, DOI: 10.1109/JETCAS.2025.3558300

Journal refIEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. 15, no. 2, pp. 231-243, June 2025

DOI:10.1109/JETCAS.2025.3558300

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑