面向异构边缘SoC上AI推理工作负载的联合硬件配置选择与映射的智能体设计空间探索
Agentic Design Space Exploration for Joint Hardware Configuration Selection and Mapping of AI Inference Workloads on Heterogeneous Edge SoCs
查看机构详情
- Purdue University(普渡大学)
- University of Chicago(芝加哥大学)
- Argonne National Laboratory(阿贡国家实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
TraceDSE提出智能体设计空间探索流程,利用执行轨迹反馈联合优化异构边缘SoC上AI推理的映射与配置,相比黑盒优化提升Pareto前沿超体积达68%并减少6-9倍硬件评估。
中文摘要 AI 辅助
现代边缘片上系统(SoC)集成了异构处理单元(PU),如CPU、GPU和NPU,每种处理单元具有不同的性能和能耗特性。在实时延迟和能耗约束下将AI推理工作负载部署到这些处理单元上,需要联合地将工作负载映射到处理单元并配置每个处理单元(例如,选择活动核心数量和运行频率)。这一联合空间呈组合式增长,使得穷举搜索不可行。先前大多数关于设计空间探索(DSE)的工作采用黑盒优化(BBO),如进化搜索,其中每次评估仅返回延迟和能耗等聚合指标。最近的LLM引导的DSE依赖于相同的稀疏反馈。我们观察到这限制了其效率:它无法提供对设计空间或设计选择为何如此表现的洞察,并且很大程度上未利用LLM的推理能力。我们提出了TraceDSE,一种智能体DSE流程,用于在异构SoC上执行AI推理的联合工作负载映射和PU配置选择。TraceDSE是一个迭代的提议者-批评者循环,由系统执行轨迹形式的更丰富反馈驱动。LLM提议者智能体生成候选映射和PU配置以供硬件评估。配备程序化轨迹分析工具的LLM批评者智能体分析轨迹以识别瓶颈并提出有针对性的改进。这一循环对每个设计点产生更深入的洞察、更高质量的决策和更有效的搜索。在Intel Meteor Lake SoC上的四个AI推理工作负载(不同复杂度的模型和多模型流水线)中,TraceDSE持续优于两种最先进的BBO工具,将Pareto前沿超体积比NSGA-II提高高达35%,比贝叶斯优化提高高达68%,同时所需的硬件评估次数减少约6-9倍。
英文摘要
Modern edge Systems-on-Chip (SoCs) integrate heterogeneous processing units (PUs) such as CPUs, GPUs, and NPUs, each with distinct performance and energy characteristics. Deploying AI inference workloads on them under real-time latency and energy constraints requires jointly mapping workloads to PUs and configuring each PU (e.g., selecting the number of active cores and the operating frequency). This joint space grows combinatorially, making exhaustive search infeasible. Most prior work on design space exploration (DSE) applies black-box optimization (BBO) such as evolutionary search, where each evaluation returns only aggregate metrics such as latency and energy. Recent LLM-guided DSE relies on the same sparse feedback. We observe that this limits its efficiency: it offers no insight into the design space or the reasons a design choice performs the way it does, and it leaves the reasoning abilities of LLMs largely unused. We present TraceDSE, an agentic DSE flow that performs joint workload mapping and PU configuration selection for AI inference on heterogeneous SoCs. TraceDSE is an iterative proposer-critic loop driven by richer feedback in the form of system execution traces. The LLM proposer agent generates candidate mappings and PU configurations for hardware evaluation. The LLM critic agent, equipped with programmatic trace-analysis tools, analyzes the traces to identify bottlenecks and suggest targeted refinements. This loop yields deeper insight into each design point, higher-quality decisions, and a more effective search. Across four AI inference workloads (models of varying complexity and a multi-model pipeline) on an Intel Meteor Lake SoC, TraceDSE consistently outperforms two state-of-the-art BBO tools, improving Pareto frontier hypervolume by up to 35% over NSGA-II and up to 68% over Bayesian optimization, while requiring ~6-9x fewer hardware evaluations.