AI 中文总结
FSGen是面向LLM的敏捷加速器生成框架,支持融合算子数据流与稀疏性,其PPA评估器精度更高,能找到功耗效率提升1.4倍或加速比达10倍的帕累托最优设计,大幅缩短设计空间探索时间,品质因数提升58倍。
AI 中文摘要
随着人工智能(AI)应用需求不断增长,大语言模型(LLM)已成为众多领域的重要 workload。如何高效生成最优AI芯片加速器设计仍是未解决且极具挑战性的问题。当前,缺乏用于高效设计空间探索(DSE)的端到端设计方法学。我们提出FSGen,这是一个基于注意力机制的LLM加速器生成敏捷框架,具备早期PPA(功耗、性能、面积)评估器。FSGen支持融合算子数据流与稀疏性,拥有多样的设计空间,与现有工作相比,能找到功耗效率提升1.4倍或加速比达10倍且PPA指标相近的设计;帕累托最优设计在各类LLM基准测试上性能显著更优,品质因数(FoM)提升58倍。由于我们的PPA评估器精度优于现有技术,且大幅缩短了DSE运行时间,设计探索速度也更快。
英文摘要
With the growing demand of artificial intelligence (AI) applications, large language models (LLMs) have become important workloads in many domains. The question of how to efficiently generate optimal AI chip accelerator designs remains unresolved and challenging. Currently, there is a lack of end-to-end design methodologies for efficient design space exploration (DSE). We propose FSGen, an agile framework for attention-based LLM accelerator generation with an early-stage PPA estimator. FSGen supports fused operator dataflows and sparsity with a diverse design space and finds designs with 1.4x better power efficiency or 10x speedup with similar PPA metrics compared to prior work. Pareto-optimal designs have much better performance over a wide range of LLM benchmarks and have 58x better figures of merit (FoM). Design exploration is also faster due to our PPA estimators, which have better accuracy than prior art and reduce DSE runtime drastically.
CommentsResearch Manuscript Published in Design Automation Conference (DAC) 63 (2026)