arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FSGen:面向大语言模型应用的具备精确功耗模型的敏捷融合与稀疏加速器生成器

FSGen: Agile Fused and Sparse Accelerator Generator with Accurate Power Model for LLM Applications

Jay Zhe-An Mok, Qijun Zhang, Zhiyao Xie

arXiv 2608.09252首次发表:更新:

AI 中文总结

FSGen是面向LLM的敏捷加速器生成框架,支持融合算子数据流与稀疏性,其PPA评估器精度更高,能找到功耗效率提升1.4倍或加速比达10倍的帕累托最优设计,大幅缩短设计空间探索时间,品质因数提升58倍。

AI 中文摘要

随着人工智能(AI)应用需求不断增长,大语言模型(LLM)已成为众多领域的重要 workload。如何高效生成最优AI芯片加速器设计仍是未解决且极具挑战性的问题。当前,缺乏用于高效设计空间探索(DSE)的端到端设计方法学。我们提出FSGen,这是一个基于注意力机制的LLM加速器生成敏捷框架,具备早期PPA(功耗、性能、面积)评估器。FSGen支持融合算子数据流与稀疏性,拥有多样的设计空间,与现有工作相比,能找到功耗效率提升1.4倍或加速比达10倍且PPA指标相近的设计;帕累托最优设计在各类LLM基准测试上性能显著更优,品质因数(FoM)提升58倍。由于我们的PPA评估器精度优于现有技术,且大幅缩短了DSE运行时间,设计探索速度也更快。

英文摘要

With the growing demand of artificial intelligence (AI) applications, large language models (LLMs) have become important workloads in many domains. The question of how to efficiently generate optimal AI chip accelerator designs remains unresolved and challenging. Currently, there is a lack of end-to-end design methodologies for efficient design space exploration (DSE). We propose FSGen, an agile framework for attention-based LLM accelerator generation with an early-stage PPA estimator. FSGen supports fused operator dataflows and sparsity with a diverse design space and finds designs with 1.4x better power efficiency or 10x speedup with similar PPA metrics compared to prior work. Pareto-optimal designs have much better performance over a wide range of LLM benchmarks and have 58x better figures of merit (FoM). Design exploration is also faster due to our PPA estimators, which have better accuracy than prior art and reduce DSE runtime drastically.

CommentsResearch Manuscript Published in Design Automation Conference (DAC) 63 (2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑