arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CIMERA:用于大语言模型推理的具有可重构精度的互连和内存计算

CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference

Yue Jiet Chong, Yimin Wang, Wei Zhang, Xuanyao Fong

arXiv 2607.13649首次发表:更新:

AI 中文总结

针对大语言模型推理能耗高的问题,提出CIMERA可重构精度推理加速器,集成计算与内存,缓解内存墙,实现精度感知执行,相比英伟达H100,处理1B和13B模型时能源效率分别提升25倍和10倍。

AI 中文摘要

大语言模型对计算和内存要求很高,给从数据中心到功率受限的边缘设备等跨平台的节能推理带来挑战。权重精度在平衡推理准确性、吞吐量和能耗方面起着关键作用,而现代大语言模型工作负载具有明显的异构性和容错性,有利于自适应精度执行。本文提出了CIMERA,一种可重构精度的大语言模型推理加速器,它集成了互连和内存计算,以缓解内存墙并实现精度感知执行。与英伟达H100相比,CIMERA在处理1B和13B模型时,能源效率分别提高了25倍和10倍。

英文摘要

LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge devices. Weight precision plays a critical role in balancing inference accuracy, throughput, and energy consumption, while modern LLM workloads exhibit pronounced heterogeneity and tolerance that favors adaptive precision execution. This paper presents CIMERA, a reconfigurable-precision LLM inference accelerator that integrates compute-in-interconnect and memory to mitigate the memory wall and enable precision-aware execution. Compared to Nvidia H100, CIMERA delivers up to $25\times$ and $10\times$ higher energy efficiency for 1B and 13B models, respectively.

CommentsAccepted to 2026 IEEE 8th International Conference on Artificial Intelligence Circuits and Systems (AICAS'26)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑