发表机构
University of Science and Technology of China; Suzhou Institute for Advanced Research, University of Science and Technology of China(中国科学技术大学; 中国科学技术大学苏州高等研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Beacon是一种报告驱动的LLM多智能体框架,用于异构多小芯片深度学习加速器的HW-DSE,可在有限迭代预算下将延迟-能耗-金钱成本的综合目标降低25.1%-93.5%。
AI 中文摘要
异构多小芯片加速器允许对各小芯片进行独立配置,以更好匹配不同算子特性并提升推理效率。然而,异构性导致模拟器评估成本高昂,限制了硬件设计空间探索(HW-DSE)可承担的迭代次数。主流数据驱动方法主要依赖最终指标和少量预定义状态,需大量搜索迭代来隐式学习输入参数与优化目标间的关系,在该场景下效果欠佳。实际中,评估器会生成关于执行时间线、资源利用率、内存访问及通信行为的详细报告。大语言模型(LLM)可将领域知识与这些报告结合,明确识别瓶颈位置、性能下降原因及参数调整方向,从而在有限迭代预算下优化每个设计决策。基于此,我们提出Beacon,一种面向异构多小芯片HW-DSE的报告驱动型LLM多智能体框架。Beacon采用分层智能体实现瓶颈定位、根本原因诊断及硬件候选生成,同时搭配分析工具包和RAG记忆以实现闭环搜索。在相同有限迭代预算下,Beacon相较于随机搜索、贝叶斯优化及强化学习,将延迟-能耗-金钱成本的综合目标降低了25.1%至93.5%。
英文摘要
Heterogeneous multi-chiplet accelerators allow chiplets to be configured independently to better match different operator characteristics and improve inference efficiency. However, heterogeneity makes simulator evaluation expensive, limiting the number of iterations affordable for hardware design space exploration (HW-DSE). Mainstream data-driven methods rely mainly on final metrics and a few predefined states, and require many search iterations to implicitly learn the relationships between input parameters and optimization objectives, making them less effective in this setting. In practice, evaluators also generate detailed reports on execution timelines, resource utilization, memory accesses, and communication behavior. Large language models (LLMs) can combine domain knowledge with these reports to explicitly identify bottleneck locations, degradation causes, and parameter adjustment directions, thereby improving each design decision under limited iteration budgets. Based on this observation, we propose Beacon, a report-driven LLM multi-agent framework for heterogeneous multi-chiplet HW-DSE. Beacon employs hierarchical agents for bottleneck localization, root-cause diagnosis, and hardware candidate generation, together with an Analysis Toolbox and RAG memory for closed-loop search. Under the same limited iteration budget, Beacon reduces the composite latency-energy-monetary-cost objective by 25.1\%--93.5\% compared with random search, Bayesian optimization, and reinforcement learning.